How to crawl a site without being a problem
September 10th 2026 · Akash Rajpurohit
Most crawlers that get blocked were not blocked for being sophisticated. They were blocked for being rude in ways their author never noticed, usually by hammering a small site with unlimited concurrency because the default HTTP client had no limits.
The useful insight is that politeness and efficiency point the same way. Nearly everything that makes you a good citizen also makes your crawl cheaper.
TLDR
- Concurrency per host is the number that matters. Global concurrency is fine; twenty simultaneous connections to one small server is not.
- Conditional requests are the biggest single win. A 304 costs both sides almost nothing.
- Back off on 429 and 5xx, exponentially and with jitter. Retrying hard turns a blip into a ban.
- Read robots.txt and the sitemap. One tells you where not to go, the other tells you exactly where to go, which saves you discovering it the expensive way.
- Crawl what changed, not everything. Most re-crawls are re-fetching pages that are byte-identical to last time.
Why per-host concurrency is the number that matters
A crawler pulling a thousand pages a minute spread across five hundred domains is invisible. The same crawler pointed at one domain is an incident.
Servers are provisioned for their normal traffic. A small documentation site might see a few requests a second. Twenty concurrent connections from one client can be several times its usual load, arriving instantly, from a single IP, with no human pattern to it. From the operator’s side that is indistinguishable from an attack, and the response is the same.
The default that will keep you out of trouble almost everywhere is one request at a time per host, with a short delay between them, and however much global concurrency you like across different hosts. If a site is large and clearly built for traffic, you can raise it. Start from politeness and earn your way up, rather than starting fast and getting cut off.
Adapt to what the server tells you
Response time is a signal. If a host normally answers in 200ms and starts taking two seconds, you are probably the reason, and continuing at the same rate makes it worse for everyone including you.
A crawler that watches latency and slows down when it rises is better than one with a fixed delay, because the right delay is not a constant. It depends on the host, the time of day, and what else that server is doing.
The same applies to status codes. A 429 is an explicit instruction. A run of 5xx usually means you have found the edge of what the origin can serve.
Handle 429 and 503 properly
The failure mode here is retrying immediately. That converts a temporary limit into a sustained overload and reliably gets you blocked.
The correct behaviour:
- If
Retry-Afteris present, honour it exactly. The server has told you the answer. - Otherwise back off exponentially: wait, then double, up to a ceiling.
- Add jitter, so that a fleet of workers hitting the same limit does not synchronise into a thundering herd on every retry.
- Give up after a bounded number of attempts and record the failure rather than looping forever.
Jitter matters more than people expect once you have more than one worker. Without it, every worker backs off by the same amount and they all return at the same instant, which is the pattern you were trying to avoid.
Conditional requests are the biggest win available
This is the part most crawlers skip, and it is worth more than every other optimisation combined.
When you fetch a page, keep the ETag and Last-Modified from the response. On the next crawl, send them back as If-None-Match and If-Modified-Since. A well-configured server answers 304 Not Modified with no body at all.
You save the transfer, you save the parsing, and the origin saves the render. On a re-crawl of a site that mostly has not changed, which is most re-crawls, this can eliminate the large majority of the real work.
Not every server implements it correctly, so you still need a fallback path. But when it works it is free, and the sites most likely to support it properly are the large ones you most want to be gentle with.
Use the map the site already published
Two files tell you most of what you need before you fetch a single content page.
robots.txt lists paths that are off limits, and often a Crawl-delay. Honour both. It also frequently points at the sitemap.
The sitemap is the higher-value one and it is routinely ignored. It lists the URLs the site considers canonical, and often lastmod dates. That is the site telling you exactly what it has and when it last changed, which lets you skip discovery crawling entirely and re-fetch only what moved.
Crawling a site by following links when it publishes a sitemap is doing unnecessary work and generating unnecessary load to arrive at a worse list.
Crawl what changed
The most common waste in production crawling is a nightly full re-crawl of a site that changes weekly.
Better strategies, in rough order of effort:
- Use
lastmodfrom the sitemap to skip anything that has not moved. - Use conditional requests so that even the pages you do fetch are cheap when unchanged.
- Learn per-section rates. A news site’s front page changes hourly; its 2019 archive does not. Crawling both at the same cadence is wrong in both directions.
- Watch feeds where they exist. RSS and Atom are still the cheapest possible change notification, and a lot of sites still publish them.
Every one of these reduces your cost and the origin’s load at the same time. That is the recurring theme: on crawling, the efficient thing and the polite thing are almost always the same thing.
Identify yourself, if you are crawling as yourself
If you are running a crawler at scale on your own behalf, a user agent that says who you are and how to reach you is worth having. It turns an anonymous load spike into a conversation. Operators who would otherwise block a mystery client will often just email and ask you to slow down.
This is not universal advice. It depends on what you are doing and on whose behalf, and there are legitimate contexts where it does not apply. But for anyone building a product that crawls the open web repeatedly, being contactable is usually in your interest.
The summary
Crawl one host at a time, watch how it responds, back off when told to, ask for what changed rather than everything, and read the two files the site published to help you. That is most of it, and it costs less than the alternative.
If you would rather not build and maintain all of that yourself, crawl does it as a managed job: per-host throttling and backoff, sitemap-aware discovery, and clean markdown out the other end at one credit a page, with failures never billed. A work email gets you 500 free credits.
[ FAQ ]
How fast can I crawl a website?
There is no universal number. Treat one request at a time per host as the default, measure the site's response times, and slow down when they rise. A site that starts responding slowly is telling you something.
Do I have to obey robots.txt?
Legally it varies by jurisdiction and situation, so this is not legal advice. Practically, obeying it is the difference between being treated as a crawler and being treated as an attack, and it costs you almost nothing.
What is the single most effective way to reduce crawl load?
Conditional requests. Send If-Modified-Since or If-None-Match and a well-configured server answers 304 with no body, so you skip both the transfer and the parsing for anything unchanged.
How should I handle a 429 or a 503?
Stop, honour Retry-After if it is present, and back off exponentially with jitter if it is not. Retrying immediately turns a temporary limit into a sustained one and gets you blocked.
Should I identify my crawler?
If you are crawling on your own behalf at scale, yes, with a way to contact you. It turns an anonymous load spike into something an operator can ask you about instead of simply blocking.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.