[ Blog ]

Notes on clean web data.

Extraction quality, feeding LLMs the web, and what we learn running the engine. Every number here is one we measured and can reproduce.

Topics

17 posts · 6 topics

Subscribe by RSS

September 25th 2026 · 4 min read

How to track competitor pricing pages

Pricing pages are designed to be read by humans and are hostile to structured extraction. What to capture, what to ignore, and how to avoid a table of noise.

September 23rd 2026 · 4 min read

How to build a chatbot over your documentation

The retrieval part is mostly solved. What decides whether a docs bot is useful is the ingestion, and that is where nearly every one of these projects goes wrong.

September 20th 2026 · 4 min read

How to find every page on a website

Sitemaps, link crawling, search operators and archives each miss something different. What each method finds, what it misses, and how to combine them.

September 16th 2026 · 5 min read

How to monitor a website for changes

Detecting that a page changed is easy. Detecting that it changed in a way you care about is the actual problem, and where most monitoring falls over.

September 14th 2026 · 5 min read

How to give ChatGPT and Claude access to a live website

Pasting a URL into a chat is not the same as giving a model access to the web. The four ways to do it properly, and when each one is the right answer.

September 10th 2026 · 5 min read

How to crawl a site without being a problem

Rate limits, robots.txt, conditional requests and the crawl etiquette that keeps you welcome. Being polite is also the cheapest way to crawl.

September 7th 2026 · 6 min read

How to make your site readable by AI agents

llms.txt, agents.md, markdown responses and the rest of the agent-readable web, with an honest account of what is a standard and what is a convention.

September 5th 2026 · 6 min read

Jina Reader alternatives for cleaner markdown

Why teams outgrow r.jina.ai, what the measured quality difference is, and which alternative fits which reason for leaving.

September 3rd 2026 · 7 min read

Firecrawl alternatives, honestly compared

What each web data API is actually good at, where Firecrawl still wins, and the measured numbers behind both claims.

September 1st 2026 · 6 min read

How to choose a web data API

The questions that actually separate web data vendors, and the afternoon-long test that answers them better than any benchmark table, ours included.

August 29th 2026 · 4 min read

How many tokens does a web page actually cost?

Raw HTML is mostly markup your model pays for and cannot use. What a page really costs in tokens, and how much of that bill is avoidable.

August 27th 2026 · 5 min read

What web scraping actually costs

The per-page price is the smallest line in the budget. What build-versus-buy really costs once maintenance, failure rates and engineering time are counted honestly.

August 25th 2026 · 7 min read

What websites cost an AI to read

We measured 40 well-known sites. The gap between the markup they send and the content they deliver runs from 6 to 1 up to 918 to 1.

August 24th 2026 · 5 min read

How to measure extraction quality yourself

Vendors publish numbers on corpora they chose. Here is how to run the same measurement on your own pages in an afternoon, including the metric worth using.

August 23rd 2026 · 5 min read

How to ground an AI agent in the live web

Search APIs hand agents a list of links. Grounding needs the page. How search-then-scrape and MCP tools give an agent something it can actually cite.

August 17th 2026 · 5 min read

Web data for RAG: what to crawl and how often

Boilerplate is the quiet tax on retrieval. How to crawl a site into a RAG index, what to skip, and how fresh the pages actually need to be.

August 12th 2026 · 5 min read

What clean web data for LLMs actually means

Clean web data is measured, not promised. What separates LLM-ready markdown from raw HTML, with the real numbers.

[ Start ]

Reading about it is one thing. Run it on your own URLs.

Every figure in these posts came from the same engine your key will hit.

Success rate

 

Median scrape

 ms