How many tokens does a web page actually cost?
August 29th 2026 · Akash Rajpurohit
Everyone building on web data eventually asks the same question at the end of the month: why is this costing so much? Usually the answer is not the model, the prompt or the retrieval strategy. It is that you are paying to tokenise navigation menus.
TLDR
- Raw HTML is mostly markup. The same Wikipedia article is 59,769 tokens as HTML and 6,831 as clean markdown, an 89 percent reduction for identical content.
- You pay for those tokens twice: once at ingestion, and again on every retrieval that pulls a bloated chunk into context.
- The waste is not evenly spread. Documentation sites and news articles are usually the worst offenders, because their templates are the heaviest part of the page.
- Cutting markup is not the same as cutting content. Headings, tables, lists and links carry meaning and must survive.
- Measure the ratio before you optimise anything else. It is usually the largest single number you can change without touching your model.
What is actually in an HTML page?
Open the source of any article you like and read it honestly. A modern page contains, roughly in order of size: the visible article, the site’s navigation, a cookie and consent banner, several analytics snippets, inline critical CSS, a footer with dozens of links to other sections, related-article widgets, social share buttons, structured data blocks, and often an inlined SVG icon set.
Exactly one of those things answers the question your user asked.
Every other item still tokenises. class="mt-4 flex items-center gap-2 rounded-lg border" is eleven or twelve tokens that mean nothing to a language model reasoning about the article’s subject. A page with two hundred such attributes has spent a couple of thousand tokens describing its own layout.
Why does this cost you twice?
The obvious cost is ingestion. If you are embedding pages for retrieval, you pay to process everything you feed in.
The less obvious cost is retrieval, and it is usually the bigger one. Boilerplate does not just inflate the document, it pollutes the chunks. A chunk that is half navigation is half wasted context every single time it is retrieved, for the entire life of the index. Worse, because site chrome repeats identically across every page, it makes unrelated pages look similar to each other, which degrades the retrieval itself.
So the same markup is charged at ingestion, charged again at every retrieval, and in between it quietly makes your search worse.
How big is the difference in practice?
Here is a real page, measured end to end rather than estimated.
| tokens | |
|---|---|
| Wikipedia article, raw HTML | 59,769 |
| Same article, clean markdown | 6,831 |
| Difference | 89% fewer |
The article is unchanged. Every heading, every list, every table and every link survives. What is gone is the part that was never the article.
The ratio varies by site. Content-heavy pages with light templates do better than the average; documentation sites and app-shell pages, where the template dwarfs the content, do dramatically worse than the average and gain the most from extraction.
What should survive extraction?
This is where most naive approaches fail. It is easy to strip a page down to plain text and call it clean. It is also a mistake, because plain text throws away structure that carries real meaning:
- Headings tell a model how the document is organised, and give chunkers sane boundaries.
- Lists signal that items are peers, not prose.
- Tables are relational data. Flattened to text, a table becomes an unreadable run of numbers.
- Links carry the destination, which is often the most useful fact on the page.
- Code blocks must keep their whitespace or they are worse than useless.
Good extraction is not “remove the tags”. It is “keep the document, remove the furniture”.
How do you check whether your pipeline is wasting tokens?
Run this on ten pages you actually ingest:
- Fetch the page and count tokens in the raw HTML, using the same tokeniser your model uses.
- Run whatever extraction you use today, and count again.
- Divide.
If your ratio is close to one, you are not extracting, you are re-encoding. If it is somewhere around a tenth, your extraction is doing real work. Then read three of the outputs by eye and check the tables and headings are still there, because a very small output can also mean you are silently dropping content, which is the more expensive failure.
That last check matters more than the ratio. A pipeline that discards half the article scores beautifully on token reduction and quietly makes your answers wrong.
Where the savings actually come from
Nothing above requires a different model or a cleverer prompt. It requires that the text you put in front of the model is the text a human would have read, and nothing else.
That is a boring engineering problem with a measurable answer, which is the best kind.
You can check it on your own pages in about a minute. Paste a URL into the free URL to markdown tool and it shows the raw HTML size, the clean markdown size and the percentage between them, with no signup. If the numbers hold up on your pages, the scrape endpoint returns the same output at one credit a page, and a work email gets you 500 free credits to run it across a real sample.
[ FAQ ]
How many tokens is a typical web page?
It varies enormously with the site, but raw HTML is usually several times larger than the content it carries. A Wikipedia article we serve is 59,769 tokens as raw HTML and 6,831 tokens as clean markdown, which is the same content for 89 percent fewer tokens.
Why is HTML so expensive in tokens?
Most of an HTML document is not the article. Inline scripts, style attributes, tracking snippets, SVG paths, navigation menus and footer link farms all tokenise, and none of it answers the user's question.
Does stripping HTML lose information the model needs?
Not if it is done well. Headings, lists, tables, links and code blocks all carry meaning and should survive. What should go is chrome: navigation, cookie banners, ads, share widgets and boilerplate that repeats on every page of the site.
Is a smaller context always better?
No, but a cheaper one usually is. The goal is not fewer tokens for their own sake, it is a higher ratio of signal to markup, so the tokens you do spend are on content the model can reason about.
How do I measure this for my own pages?
Fetch a page, count the tokens in the raw HTML, then count them again after extraction with the same tokeniser your model uses. The ratio between the two is the part of your bill that was never content.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.