Web data for RAG: what to crawl and how often
August 17th 2026 · Akash Rajpurohit
Most RAG pipelines are built on the assumption that getting the text is the easy part. It is not, and the failures are quiet: retrieval returns a cookie banner, an answer cites a nav link, and the index costs more than it should. This post covers what to crawl into a retrieval index, what to leave out, and how often to go back.
TLDR
- Chunk clean markdown, not raw HTML. A Wikipedia article is 59,769 tokens as HTML and 6,831 tokens as clean markdown, so an HTML chunk carries roughly a tenth of the meaning per token.
- Boilerplate is embedded once per page. A 300 page site embeds its own footer 300 times, and those chunks match queries they should never match.
- Crawl, do not write a link loop. One call takes a domain, depth and path rules, and calls a webhook when it finishes.
- Set freshness to how fast the content moves, not to a habit. Docs weekly, reference rarely, news in hours.
- Skip pages the response tells you are thin. Every response carries a confidence figure, so a bad extraction can be dropped before it reaches the index.
Why does boilerplate hurt retrieval?
Because it is repeated, and repetition is exactly what embeddings are good at finding. Navigation, the cookie notice, the newsletter box and the footer appear on every page of a site. Split a page into chunks without removing them and each one becomes an embedded vector like any other.
Two things go wrong. The first is cost: you pay to embed and store text nobody will ever want returned. The second is worse, because it is invisible. Those chunks are textually similar to every other page on the same site, so they surface for queries that have nothing to do with them, and they push a genuinely relevant passage out of your top-k.
The fix is not better chunking. It is not putting the boilerplate in the index in the first place.
Should I chunk HTML or markdown?
Markdown, and the reason is arithmetic. Splitters work in tokens, and raw HTML spends most of its tokens on markup that carries no meaning for retrieval.
| Raw HTML | Clean markdown | |
|---|---|---|
| Characters | 239,076 | 27,326 |
| Tokens | 59,769 | 6,831 |
That is the same Wikipedia article, at 11 percent of the token count. A 1,000 token chunk of markdown holds roughly ten times the meaning of a 1,000 token chunk of HTML, so your top-k returns ten times the substance for the same context budget.
Structure matters too. A model reasons better over a heading followed by a table than over the same words flattened into a paragraph, and a splitter can use headings as natural boundaries. Extraction that strips markup but also strips structure has solved half the problem.
How do I crawl a site instead of scraping one page?
Send the domain and the rules, not a loop. A crawl runs as an asynchronous job with depth, path and subdomain controls, and calls your webhook when it is done.
curl -X POST https://api.hydrafetch.com/v1/web/crawl \
-H "X-API-Key: $HYDRAFETCH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.example.com",
"limit": 500,
"maxDepth": 3,
"excludePaths": ["/changelog/*"],
"webhook": { "url": "https://yourapp.com/hooks/hydrafetch" }
}'
excludePaths is worth more thought than it usually gets. Changelogs, tag pages, paginated archives and author pages tend to be high volume and low value for retrieval. Excluding them up front costs nothing and keeps both your bill and your index smaller.
Every page comes back in the same envelope, so ingestion is one code path:
for page in pages:
data = page["data"]
if data["quality"]["confidence"] < 0.5:
continue
chunks = splitter.split_text(data["markdown"])
index.upsert([
{
"text": chunk,
"source": data["finalUrl"],
"title": data["metadata"].get("title"),
}
for chunk in chunks
])
Two details in that loop earn their place. Storing finalUrl rather than the URL you requested means redirects do not break your citations. Dropping low-confidence pages means a bad extraction is skipped rather than quietly poisoning retrieval, which is the failure mode you will never notice from the inside.
How fresh does the data need to be?
Freshness is a per-source decision, and most pipelines get it wrong in both directions at once: re-crawling reference material nightly, and letting a pricing page go stale for a month.
| Content | Changes | Worth re-crawling |
|---|---|---|
| Documentation, pricing, policies | Weeks | Weekly |
| Product and catalog pages | Days to weeks | Weekly |
| Blogs and news | Hours to days | Daily or faster |
| Reference and archives | Rarely | On demand |
You can also just say how fresh is fresh enough on the request. Recently fetched pages are served from a short-lived cache and refreshed when they go stale, and every response states whether it came from cache, so you can measure your own hit rate instead of guessing.
What does it cost to keep an index current?
One credit per page fetched, whatever that page took to deliver, and nothing for pages that fail. That makes the arithmetic simple enough to do in your head: a 500 page documentation site costs 500 credits per full crawl, so weekly refreshes are about 2,000 credits a month.
The part that is easy to miss is that re-processing is not re-fetching. Raw HTML is stored immutably, so when you change your chunking strategy or want a different format from pages you already paid for, you are not paying to crawl the site again. In a pipeline you will rebuild several times before you are happy, and that difference adds up faster than the crawl itself.
Does extraction quality actually vary between vendors?
Yes, and it is measurable, which means you do not have to take anyone’s word for it. The standard is word-level F1 against pages where humans have marked which words are main content. It punishes both keeping junk and dropping content.
Our engine scores 0.859 word-level F1 on the held-out split of a 2,008 page human-labeled corpus. Trafilatura, the strong open-source baseline, scores 0.857 on the same pages, so on one overall number the two of us are level. Jina’s own ReaderLM-v2 model scores 0.741 on the development split.
The number to ask any vendor for is the F1 and the denominator. We wrote about how to run that check yourself in what clean web data for LLMs actually means, and the RAG use case page has the shorter version of this pipeline.
Where to start
Crawl one section of one site, not the whole domain. Look at fifty chunks by hand and ask whether you would want any of them returned for a real query. That hour tells you more about your retrieval quality than any benchmark, and it is usually where people discover their index is a third boilerplate.
Then set the crawl to run on the cadence the content actually moves at, and let the confidence figure drop the pages that came back thin.
[ FAQ ]
What is the best way to get web data into a RAG pipeline?
Crawl the site once into clean markdown, chunk that rather than raw HTML, and store the source URL with every chunk. Crawling gives you the whole site in one job instead of you writing a link discovery loop.
Why does boilerplate hurt retrieval quality?
Navigation, cookie notices and footers repeat on every page, so they get embedded once per page and compete with real content at query time. On a 300 page site you have embedded the same footer 300 times.
How often should I re-crawl a site for RAG?
Match the crawl to how fast the content actually changes. Documentation and pricing pages are worth re-crawling weekly, reference and archive content far less often, and news within hours if it matters to you.
Should I chunk HTML or markdown?
Markdown. Splitters count tokens, and raw HTML spends most of its tokens on markup, so an HTML chunk carries a fraction of the meaning of a markdown chunk the same size.
How much does crawling a site cost?
One credit per page fetched, whatever the page took to deliver, and pages that fail cost nothing.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.