[ Blog ]
Notes on clean web data.
Extraction quality, feeding LLMs the web, and what we learn running the engine. Every number here is one we measured and can reproduce.
September 25th 2026 · 4 min read
How to track competitor pricing pages
Pricing pages are designed to be read by humans and are hostile to structured extraction. What to capture, what to ignore, and how to avoid a table of noise.
September 23rd 2026 · 4 min read
How to build a chatbot over your documentation
The retrieval part is mostly solved. What decides whether a docs bot is useful is the ingestion, and that is where nearly every one of these projects goes wrong.
September 20th 2026 · 4 min read
How to find every page on a website
Sitemaps, link crawling, search operators and archives each miss something different. What each method finds, what it misses, and how to combine them.
September 16th 2026 · 5 min read
How to monitor a website for changes
Detecting that a page changed is easy. Detecting that it changed in a way you care about is the actual problem, and where most monitoring falls over.
September 14th 2026 · 5 min read
How to give ChatGPT and Claude access to a live website
Pasting a URL into a chat is not the same as giving a model access to the web. The four ways to do it properly, and when each one is the right answer.
September 10th 2026 · 5 min read
How to crawl a site without being a problem
Rate limits, robots.txt, conditional requests and the crawl etiquette that keeps you welcome. Being polite is also the cheapest way to crawl.
September 7th 2026 · 6 min read
How to make your site readable by AI agents
llms.txt, agents.md, markdown responses and the rest of the agent-readable web, with an honest account of what is a standard and what is a convention.
September 5th 2026 · 6 min read
Jina Reader alternatives for cleaner markdown
Why teams outgrow r.jina.ai, what the measured quality difference is, and which alternative fits which reason for leaving.
September 3rd 2026 · 7 min read
Firecrawl alternatives, honestly compared
What each web data API is actually good at, where Firecrawl still wins, and the measured numbers behind both claims.
September 1st 2026 · 6 min read
How to choose a web data API
The questions that actually separate web data vendors, and the afternoon-long test that answers them better than any benchmark table, ours included.
August 29th 2026 · 4 min read
How many tokens does a web page actually cost?
Raw HTML is mostly markup your model pays for and cannot use. What a page really costs in tokens, and how much of that bill is avoidable.
August 27th 2026 · 5 min read
What web scraping actually costs
The per-page price is the smallest line in the budget. What build-versus-buy really costs once maintenance, failure rates and engineering time are counted honestly.
August 25th 2026 · 7 min read
What websites cost an AI to read
We measured 40 well-known sites. The gap between the markup they send and the content they deliver runs from 6 to 1 up to 918 to 1.
August 24th 2026 · 5 min read
How to measure extraction quality yourself
Vendors publish numbers on corpora they chose. Here is how to run the same measurement on your own pages in an afternoon, including the metric worth using.
August 23rd 2026 · 5 min read
How to ground an AI agent in the live web
Search APIs hand agents a list of links. Grounding needs the page. How search-then-scrape and MCP tools give an agent something it can actually cite.
August 17th 2026 · 5 min read
Web data for RAG: what to crawl and how often
Boilerplate is the quiet tax on retrieval. How to crawl a site into a RAG index, what to skip, and how fresh the pages actually need to be.
August 12th 2026 · 5 min read
What clean web data for LLMs actually means
Clean web data is measured, not promised. What separates LLM-ready markdown from raw HTML, with the real numbers.
[ Start ]
Reading about it is one thing. Run it on your own URLs.
Every figure in these posts came from the same engine your key will hit.
Success rate
Median scrape
ms