---
title: "Every extractor is good at articles | Hydrafetch"
url: https://hydrafetch.com/reports/extraction-benchmark/
description: "Word-level F1 for seven content extractors across 1,497 labelled pages. On articles they agree. On everything else they do not."
---

[hydrafetch](https://hydrafetch.com/)[Book a call](https://hydrafetch.com/demo/)[Get started](https://app.hydrafetch.com/)

Endpoints

[ScrapeClean markdown from one URL](https://docs.hydrafetch.com/endpoints/scrape)[MapEvery URL on a site, one call](https://docs.hydrafetch.com/endpoints/map)[CrawlWhole sites as one job](https://docs.hydrafetch.com/endpoints/crawl)[SearchRanked results, already fetched](https://docs.hydrafetch.com/endpoints/search)[ExtractTyped records from a schema](https://docs.hydrafetch.com/endpoints/extract)[BrandA company from its domain](https://docs.hydrafetch.com/endpoints/brand)
Concepts

[FormatsMarkdown, structured, links](https://docs.hydrafetch.com/concepts/formats)[QualityConfidence on every response](https://docs.hydrafetch.com/concepts/quality)[CreditsOne a page, failures free](https://docs.hydrafetch.com/concepts/credits)[CachingFresh or cached, your call](https://docs.hydrafetch.com/concepts/caching)

[Try it Run any endpoint on a URL of your own, no key](https://hydrafetch.com/try/)
Web data for AI

[Ground RAG in fresh contentCrawl a site on a schedule and pipe clean Markdown into your embeddings.](https://hydrafetch.com/use-cases/rag/)[Give an agent the live webSearch mid-answer and get results already fetched, cleaned and citable.](https://hydrafetch.com/use-cases/agents/)[Turn listings into a datasetPoint a JSON schema at a directory and get typed rows back.](https://hydrafetch.com/use-cases/structured-extraction/)[Watch pages for changesRe-run a set on a schedule and diff what came back.](https://hydrafetch.com/use-cases/change-monitoring/)[Track prices and competitorsMap a catalogue, pull every page, and extract the fields that move.](https://hydrafetch.com/use-cases/competitor-intelligence/)
Company and brand data

[Enrich a company from its domainOne call returns the name, industry, assets, palette and socials.](https://hydrafetch.com/use-cases/company-enrichment/)[Theme your app per tenantResolve a domain to a full design system and restyle your UI at runtime.](https://hydrafetch.com/use-cases/white-label-theming/)[Autofill onboardingTurn a work email domain into a filled-in company form.](https://hydrafetch.com/use-cases/onboarding-autofill/)

[All use cases Each one worked end to end, on real captures](https://hydrafetch.com/use-cases/)
[DocsEndpoints, formats and limits](https://docs.hydrafetch.com/)[BlogHow the engine is measured](https://hydrafetch.com/blog/)[Extraction benchmarkWord-level F1 against the field](https://hydrafetch.com/reports/extraction-benchmark/)[Token indexWhat a page costs once it is clean](https://hydrafetch.com/reports/token-index/)
[Talk to a human Thirty minutes, and bring your hardest URL](https://hydrafetch.com/demo/)

Measured August 23rd 2026

# Every extractor is good at articles

Content extractors are judged on benchmarks made mostly of articles, and on articles they all work. We scored seven of them across 1,497 labelled pages of seven kinds. On articles the field is separated by 10.9 points. On service pages it is 51.1.

1,497 labelled pages1,296 domainsOpen corpus and method

The same extractor, two page typesreadability, the one inside most reader modes, keeps **82.4%** of an article and **23.3%** of a listing page.

Narrowest spread

10.9

articles, 53% of the corpus

Widest spread

51.1

service pages, 165 of them

Extractors scored

7

on identical pages and labels

## Where the tools actually differ

Sorted by how far apart the field is. Articles and documentation sit at the top because everything handles them. The pages an agent hits when it leaves a blog are at the bottom.

Each line runs from the weakest extractor on that page type to the strongest. The filled dot is Hydrafetch. The number on the right is the distance between them, in points of F1.

## Overall, across every page type

One number each, weighted by the corpus. Useful for a ranking and misleading on its own, since more than half of these pages are articles. We place second: rs-trafilatura is ahead of us here and on every page type below, and we would rather you read that from us than find it yourself.

- rs-trafilatura

84.7

- Hydrafetch

82.9

- trafilatura

81.3

- trafilatura (recall)

78.8

- resiliparse

77.1

- jusText

69.1

- readability

61.7

hydrafetch.comWord-level F1 against the labelled main content, as a percentage. Higher is better.

## What the numbers say

The easy case

An article has one column of prose, a headline and an author, and every extractor was built by someone reading one. 53% of this corpus is articles, so an overall score is mostly a score on the easy case.

What agents actually read

Not articles. Pricing tables, product pages, search results and forum threads, and those are where the field falls apart: 51.1 points between weakest and strongest on service pages, against 10.9 on articles.

The honest rebuttal

Which we would raise ourselves: if you only ever read blog posts and documentation, any of these will do, and this is a solved problem for you.

## How we measured

The corpus

1,497 human-labelled pages across 1,296 domains and seven page types, scored with WCXB's own word-level F1 against the labelled main content. It ships its saved HTML, so every run is deterministic, touches no network and involves no model. Nothing was dropped: every extractor saw every page.

The engines

We ran our live extraction path rather than a copy, so this page moves when the product does. Baselines are the open extractors at their defaults, on the versions current when this ran: trafilatura 2.2.0, resiliparse 1.0.9, jusText 3.0.2, readability 0.8.4.1. Those are pinned in the dataset, because a baseline that quietly ages is the easiest way for a benchmark to flatter its author.

Whose observation this is

Not ours. WCXB's authors built the corpus because article-only benchmarks hide exactly this, and they say so in its documentation. What is ours is the measurement, over these seven extractors, with our own scored the same way as the rest.

Limits

This is one benchmark, and a corpus that is half articles flatters tools tuned for articles. Our own rs-trafilatura figure is the softest number here: it is the one extractor we could not pin to a version, and its source is no longer resolving, so we cannot tell you how to rerun that column. WCXB's published leaderboard puts it at 0.859 on this split, a little above the 0.8471 we measured.

Check it yourself

The harness and dataset are on GitHub, MIT licensed, at [Hydrafetch/extraction-benchmark](https://github.com/Hydrafetch/extraction-benchmark). It runs the five open extractors over the same corpus and prints the same table, so every baseline number here is one command from being checked.

What we cannot prove to you

Our own column. Our API takes a URL and the corpus is saved HTML, so we score it in process against the same labels, with the same metric, over the same files. That is a weaker guarantee than the rest of the table, and we would rather say so than imply otherwise. We build a web data API, and one of these columns is ours.

[ Start ]

## Clean web data is one call away.

250 free credits, no card required. Failures are never billed.

[Get API key](https://app.hydrafetch.com/)[Read the docs](https://docs.hydrafetch.com/)
