What websites cost an AI to read
August 25th 2026 · Akash Rajpurohit
A browser reads your website once and throws the markup away. A model pays for every token of it.
That is a new cost that nobody has a budget line for. We measured 40 well-known sites to see how big it is, and the answer is that it varies by a factor of 160 depending on how the site was built.
The full ranking is in the token index. This post is what is in it.
TLDR
- The spread is 6:1 to 918:1 between markup sent and content delivered, across 40 sites measured the same way.
- The worst offenders are developer tools. Figma at 918:1, redis.io at 750:1, Atlassian at 580:1, Vercel at 440:1.
- The leanest are the plainest. postgresql.org at 6:1, Hacker News and kubernetes.io at 11:1.
- The frameworks’ own sites match their own philosophies. astro.build, svelte.dev and react.dev all land lean; nextjs.org is the outlier at 148:1 and 42% inline script.
- The cause is inlining, not one specific asset. Script on some, icon sprites on others, stylesheets on a third group. The two leanest sites inline nothing at all.
- This is invisible in every existing budget. It is not bandwidth, it is not Core Web Vitals, and a browser never pays it.
What we measured
One homepage per site, fetched once on August 23rd 2026. We counted the HTML as delivered, then counted what survived extraction, and divided. Tokens are counted at four characters each.
Any capture our engine flags as incomplete or challenged is dropped rather than published, so a site that fails to measure is absent instead of wrong. All 40 in this run came back clean.
The list is chosen for names you will recognise. It is not a sample of the web, and the median across it is a fact about those 40 sites rather than an estimate of anything wider.
The results
| site | markup sent | content | ratio |
|---|---|---|---|
| figma.com | 419,615 | 457 | 918:1 |
| redis.io | 436,034 | 581 | 750:1 |
| atlassian.com | 620,434 | 1,070 | 580:1 |
| vercel.com | 150,997 | 343 | 440:1 |
| nextjs.org | 81,711 | 554 | 148:1 |
| tailwindcss.com | 225,146 | 1,548 | 145:1 |
| stripe.com | 160,176 | 2,219 | 72:1 |
| stackoverflow.com | 110,036 | 1,399 | 79:1 |
| astro.build | 78,341 | 1,142 | 69:1 |
| svelte.dev | 22,443 | 534 | 42:1 |
| react.dev | 68,107 | 1,952 | 35:1 |
| kubernetes.io | 10,585 | 947 | 11:1 |
| news.ycombinator.com | 8,607 | 817 | 10:1 |
| python.org | 13,210 | 1,022 | 13:1 |
| postgresql.org | 5,950 | 1,045 | 6:1 |
The pattern is uncomfortable if you build for a living. The sites that developers make for other developers are the expensive ones, and several of the companies at the top sell products whose entire value proposition is web craft.
The sites that win are the ones nobody would call well designed. postgresql.org sends six tokens of markup per token of content. It beats every venture-funded developer tool in the list by more than an order of magnitude.
The framework sites are the most interesting rows, because each one is a demonstration of what its authors believe. Astro sells zero JavaScript by default, and astro.build is 1% inline script at 69:1. react.dev comes in at 35:1 and svelte.dev at 42:1, both leaner than most of the products built with them. nextjs.org is the outlier at 148:1 and 42% inline script. Nobody is being a hypocrite here. The defaults each project ships are simply visible in its own homepage, which is the fairest test of a framework anyone could ask for.
Why is it so expensive?
We measured what each document is made of, as a share of its characters. There is no single culprit, which surprised us.
| site | inline script | inline style | inline SVG | base64 |
|---|---|---|---|---|
| figma.com | 83% | 8% | 2% | 4% |
| redis.io | 82% | 0% | 3% | 46% |
| tailwindcss.com | 54% | 0% | 23% | 1% |
| nextjs.org | 42% | 0% | 24% | 1% |
| supabase.com | 8% | 0% | 82% | 0% |
| cloudflare.com | 3% | 1% | 71% | 0% |
| wikipedia.org | 2% | 51% | 9% | 0% |
| svelte.dev | 27% | 0% | 9% | 4% |
| astro.build | 1% | 1% | 61% | 0% |
| react.dev | 2% | 0% | 49% | 0% |
| news.ycombinator.com | 0% | 0% | 0% | 0% |
| postgresql.org | 0% | 0% | 0% | 0% |
Those shares are of the whole document and overlap, since base64 usually sits inside a script or style block.
Inline script is the story on some of them. Figma and redis.io are more than four fifths script, most of it not code but state: the data a framework serialises into the page so the client can rebuild what the server already rendered. On redis.io nearly half of a 1.7MB homepage is base64 sitting in the markup.
But it does not generalise. Supabase is 82% inline SVG and only 8% script. Cloudflare is 71% SVG. Wikipedia is 51% inline style. Across the list the heaviest ten average 34% script against 11% for the leanest ten, which is a real difference and nowhere near an explanation.
The thing the two leanest sites have in common is simpler than any of that. postgresql.org and Hacker News inline nothing. Zero on every measure above. Everyone else has committed to inlining something, whether that is framework state, an icon sprite or a stylesheet, and pays for it once per agent that reads the page.
Does this actually matter?
It matters in exactly one case, and that case is common.
If you feed raw HTML to a model, this ratio is your bill. That is what a naive pipeline does, and a lot of pipelines are naive: fetch the page, drop it in the context, ask a question. At 918:1 you are paying for 917 tokens of framework state for every token of content, and it also crowds out the context window you wanted for something else.
If you already clean pages before the model sees them, the ratio is the size of a problem you have solved rather than one you have. That is a fair objection and it is worth stating plainly.
What it never means is that the site is slow or badly built for humans. A browser is not billed by the token. Nothing in this measurement says anything about page speed, Core Web Vitals, or the experience of a person visiting. This is a cost that only exists when the reader is a machine, which is why nobody has been tracking it.
The reason it is worth tracking now is that the share of your readers who are machines is going up, and none of them are looking at your design.
What to do about it
Three things, cheapest first.
Publish an llms.txt. A short file pointing at the pages you want read, so an agent does not have to discover your site by crawling the expensive parts. It takes minutes and it is the highest-value thing on this list. Our llms.txt generator will draft one from your sitemap.
Serve your content in the initial HTML. If the words only exist after the framework runs, every reader that is not a browser either misses them or pays a lot more to get them.
Stop inlining. Serialised state, icon sprites and stylesheets in the document are what separates the top of the table from the bottom. Fetch them instead of embedding them, and they are cached once rather than paid for on every read.
None of this requires rebuilding anything. The gap between postgresql.org and the bottom of that table is not effort, it is a set of defaults.
Check your own site
You can run the same measurement on any page with our URL to markdown tool. It shows the tokens before and after alongside the word count, which is the part that matters: a page that returns nothing would score perfectly on markup reduction alone.
If you want to see how a page reads to an agent more broadly, the AI readability audit covers whether your content survives without JavaScript, whether assistant crawlers are allowed, and whether you have an llms.txt.
The full ranking, the method and the caveats are in the token index.
[ FAQ ]
How many tokens does a web page cost?
It depends entirely on the site. Across 40 well-known sites we measured, one homepage ranged from about 6,000 tokens of HTML to about 620,000. The content inside them ranged from 168 words to a few thousand, so the ratio of markup to content ran from 6 to 1 up to 918 to 1.
Why is my website expensive for an AI to read?
Inlining, though which asset varies. Some heavy sites are more than 80% inline script carrying framework state; others are 70 to 80% inline SVG icon sprites, and one is half inline stylesheet. The two leanest sites we measured inline nothing at all. None of it is content, and all of it is tokens if you feed the raw HTML to a model.
Does this matter if I strip the HTML first?
Then the ratio is the size of the problem you already solved. It matters if you are feeding raw pages to a model, which naive pipelines do, and it matters for how much work a crawler has to do before your content is usable. It does not matter for a browser, which renders the markup and discards it.
How do I make my site cheaper for agents to read?
Serve the content in the initial HTML rather than assembling it after load, keep data payloads out of the document, and publish an llms.txt pointing at the pages you want read. The last one costs nothing and does the most.
Is a smaller page always better?
No. A page that returns nothing scores perfectly on every size metric. The number that matters is markup relative to content that survived, which is why the index reports both and not just a reduction percentage.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.