How to measure extraction quality yourself
August 24th 2026 · Akash Rajpurohit
Every extraction vendor publishes a quality number, ours included. All of them were produced on a corpus the vendor selected. That does not make them dishonest, it makes them irrelevant to your decision, because your pages are not the average of the web.
The good news is that measuring this yourself is genuinely an afternoon of work, and it settles the question properly.
TLDR
- Word-level F1 against human-labelled pages is the standard metric. Use it.
- Twenty to thirty pages is enough to separate good from bad. You are not writing a paper.
- Choose the sample to match your mix, deliberately including the awkward cases.
- Precision and recall fail differently. Decide which one hurts you more before you look at the scores.
- Length ratios and token counts are not quality measures and can be gamed in both directions.
The metric
Treat extraction as a retrieval problem over the words of the page.
- Precision: of the words the extractor returned, what fraction should have been there
- Recall: of the words that should have been there, what fraction did it return
- F1: the harmonic mean, which punishes lopsidedness
Both directions matter and they fail differently. An extractor that returns the whole page scores perfect recall and terrible precision. One that returns the first paragraph scores well on precision and badly on recall. F1 refuses both.
Word-level, not character-level, because character-level scoring rewards whitespace and punctuation agreement that nobody cares about.
Building the sample
This is the step that determines whether the exercise is worth anything.
Take pages you actually process. Not a public benchmark, not the vendor’s examples. If you ingest documentation, use documentation. If you ingest small business sites, use those, because they break naive extractors far more often than well-built sites do.
Include the awkward ones on purpose:
- A long article, to catch truncation
- A page with a real data table
- A page behind a cookie or consent wall
- A page that is mostly JavaScript-rendered
- A page from a site with an unusual template
- Something short, like a product page, where boilerplate outweighs content
Twenty to thirty of these is plenty. You are trying to tell good from bad, not to publish.
Labelling without losing a day
Labelling sounds heavy and is not, if you do it the pragmatic way.
For each page, save the text a human would consider the content: the article, the documentation, the product description. Include headings, list items and table cells. Exclude navigation, cookie notices, share widgets, related links, and the footer.
Copying from a reader view and cleaning it is usually faster than any tooling. Fifteen pages an hour is a normal pace, so a sample of thirty is about two hours.
The one rule that matters: decide your edge cases once and apply them consistently. Are image captions content? Is the author byline? Either answer is fine. Changing your mind halfway through is what makes a labelled set worthless.
Scoring it
def f1(predicted: str, truth: str) -> float:
from collections import Counter
p, t = Counter(predicted.lower().split()), Counter(truth.lower().split())
overlap = sum((p & t).values())
if not overlap:
return 0.0
precision, recall = overlap / sum(p.values()), overlap / sum(t.values())
return 2 * precision * recall / (precision + recall)
Run every page through each extractor you are comparing, score each, and report the mean along with precision and recall separately. The split matters: two extractors with the same F1 can be wrong in opposite ways, and which one you want depends on your use case.
Read the worst ones
The number tells you where you stand. The failures tell you why, and that is the part that changes your decision.
Sort by score and read the bottom five. You will usually find one of a few things:
- Truncation. The article stops partway, which shows as good precision and poor recall.
- Boilerplate included. Navigation survived, which shows as good recall and poor precision.
- A wall. The extractor faithfully returned a consent page. This is the important one, because it will score terribly and, in production, would have looked like a success.
- Structure lost. The content is all there as prose, but the table became a run of numbers.
That last one deserves attention because F1 does not catch it. Word overlap is identical whether or not the table survived as a table. If structure matters to you, check it by eye separately.
Decide what you care about before you look
Worth settling in advance, so the result does not get rationalised afterwards:
For retrieval, recall usually dominates. A paragraph the extractor dropped is a question your system can never answer, and no amount of reranking recovers it.
For summarisation, precision usually dominates. Boilerplate that survives extraction ends up in the summary, and a summary mentioning the cookie policy is obviously broken to a user.
For structured extraction, neither is quite right. What matters is whether the specific fields you want survived, so score those directly instead.
What to do with the answer
If a vendor scores well on your sample, you have a real reason to trust it beyond the marketing. If it scores badly, you have specific pages to show them, which is a much better conversation than “the quality seems off”.
And if you are building your own, this is the harness that tells you whether a change helped. Extraction changes are exactly the kind that feel better and measure worse.
For reference on our own numbers: we score 0.859 word-F1 on the held-out split of a 2,008-page human-labelled corpus, and re-run it against the live engine before any extraction change ships. That is a number about our corpus, though, which is the whole point of this post. Run it on yours.
You can start with a single page and no signup in the URL to markdown tool, or take a work email’s 500 free credits and run the full sample through scrape.
[ FAQ ]
How do you measure web extraction quality?
Word-level F1 against pages a human has labelled. You mark what the main content is, compare the extractor's output word by word, and score precision and recall together. It is the standard used in the academic literature and it is reproducible.
Why not just compare output length?
Because length is trivially gamed in both directions. An extractor that drops half the article and one that keeps the navigation can produce identical lengths, and only one of them is wrong in a way you would notice.
How many pages do I need to label?
Fewer than you think. Twenty to thirty pages chosen to represent your real mix will separate a good extractor from a bad one clearly. Hundreds are for detecting small differences between two good ones.
What matters more, precision or recall?
It depends on what you are building. For retrieval, recall usually matters more because a missing paragraph is a question you cannot answer. For summarisation, precision matters more because included boilerplate ends up in the summary.
Does a high score mean the extractor is good for me?
Only if the corpus resembles your pages. A score on news articles predicts very little about ecommerce listings or documentation, which is why running it on your own sample is the point.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.