All posts

How to choose a web data API

September 1st 2026 · Akash Rajpurohit

Every web data vendor’s landing page says the same four things: fast, reliable, clean output, works on JavaScript sites. None of that is falsifiable from the outside, which is why comparing them by reading marketing copy does not work.

Here are the questions that do separate them, and a test you can run in an afternoon that beats every benchmark table, including ours.

TLDR

  • Test on your own URLs. Vendor benchmarks are chosen by vendors. Your ten worst pages tell you more than any published number.
  • Ask what happens when a page fails, and whether you are billed for it. This one question sorts vendors quickly.
  • Look for a quality signal on the response. Without one you cannot detect the failure that matters: a 200 that contains a consent wall.
  • Watch for multiplier pricing. If rendering, stealth or proxies cost extra, your bill depends on facts about pages you cannot know in advance.
  • Check the ending of a long article. Silent truncation is common and invisible until it corrupts an answer.

Test on your own URLs, not theirs

The single most important thing: every vendor’s published numbers were produced on a corpus that vendor selected.

That is not necessarily dishonest, it is just useless to you. Your workload is not the average of the web. It might be ecommerce product pages, or academic PDFs, or small business sites built on page builders, or documentation. Performance on one of those predicts very little about the others.

So build a list of ten URLs that represent what you actually do, and deliberately include the awkward ones:

  • One long article, to test whether the ending survives.
  • One page that is heavily JavaScript-driven.
  • One page with a cookie or consent wall.
  • One page with a real data table.
  • One page from a small site with an unusual template, because these break naive extractors far more often than famous sites do.

Run those ten through every vendor on your shortlist. This takes an afternoon and settles most of the argument.

Ask what happens when a page fails

Failure is normal in web data. Sites go down, pages are removed, some hosts refuse you. What matters is what the API does about it, and there are three good questions:

Do I get told, clearly? A failure should be distinguishable from an empty page. If both come back as a 200 with no content, you cannot build reliable ingestion on it.

Am I billed? Being charged for a page that returned a CAPTCHA means paying for someone else’s anti-bot vendor. Our position is that failures cost nothing, and we think that is the only defensible one, but the important thing is that you know the answer before you commit.

Can I detect it programmatically? You will not be reading responses by hand at volume. There must be something on the response your code can branch on.

The failure mode that costs the most

If you take one thing from this: the expensive failure is not an error, it is a success that is wrong.

A request returns 200. The body is well-formed HTML. Extraction runs cleanly and produces tidy markdown. The markdown says “Please verify you are human” or “We use cookies to improve your experience”.

Nothing in your pipeline notices. It embeds, indexes, and sits there until a user asks a question and gets a confident answer built on a consent banner.

This is why a quality signal on every response matters more than almost any other feature. Something that says: here is how confident the extraction is, here is whether the bulk of the page came through, here is whether we think a wall stood in the way. Without those, your only defence is reading output by hand, which does not scale past your first week.

Ask every vendor on your list how you would catch this. The answers are revealing.

Watch for pricing that multiplies

Two vendors quoting “one credit a page” can produce bills that differ by five times, because of what happens to that credit.

The common model is pricing the mechanism: a base rate for a simple fetch, then multipliers for JavaScript rendering, for stealth, for premium proxies, sometimes stacking. The headline number is attractive and the real bill lands several times higher.

The problem is not the price, it is the predictability. You cannot know in advance which of your URLs will need rendering. That fact is a property of the site, and it changes when the site changes, without telling you. Under multiplier pricing your costs move for reasons entirely outside your control or knowledge.

The alternative is flat per-page pricing, where the vendor absorbs that variance. We price this way deliberately: a page costs one credit whatever it took to fetch, and how hard it was is our problem. Additional product costs more, so model-backed extraction is priced above a plain fetch, because you chose it and it has real marginal cost. But mechanism is included.

Whichever model you pick, price your actual workload rather than the headline. Take your ten URLs, work out what each vendor would charge, and compare that number.

Read the output, do not just count it

Token counts and size reductions are easy to publish and easy to game, because the cheapest way to make output small is to throw content away.

So read three of the outputs properly:

  • Does the long article still end where it ended? Truncation is common and silent.
  • Did the table survive as a table? Flattened into prose, a table is worse than absent, because it looks like data and is not.
  • Are the headings intact? They are how your chunker finds boundaries.
  • Is the navigation gone? If site chrome is still there, you are paying to embed a menu on every page.

A vendor scoring beautifully on size reduction while dropping the last third of every article is the worst of both worlds, and only reading catches it.

Should you build it yourself?

Sometimes yes. If you need ten known sites with stable markup, hand-written extractors will beat a general-purpose API and cost nothing to run.

The calculation changes with breadth and time. The expensive part was never the first version, it is the long tail: thousands of sites with unusual templates, sites that redesign without warning, pages that only exist after JavaScript runs, hosts that decide to challenge you. That is a maintenance stream forever, and it does not become more interesting over time.

The honest split: few sites, stable, full control needed, build it. Many sites, changing, and the extraction is not your product, buy it.

The afternoon test

To summarise it as a checklist:

  1. Ten of your own URLs, including the awkward ones.
  2. Run them through each vendor.
  3. Read the outputs, checking endings, tables and headings.
  4. Break one deliberately, with a dead URL or a walled page, and see what comes back and whether you are billed.
  5. Price your real workload, not the headline.

That afternoon will tell you more than every comparison page on the internet, this one included.

If you want to run it against us, the scrape endpoint returns clean markdown with a quality object on every response, at one credit a page with failures never billed. You can try a single URL with no signup in the URL to markdown tool, and a work email gets you 500 free credits to run the full ten.

[ FAQ ]

What should I look for in a web scraping API?

Extraction quality on your own URLs, an honest billing model for failures, a quality signal on every response so you can catch silent failures, and predictable pricing that does not multiply for pages that happen to be harder to fetch.

How do I compare extraction quality between vendors?

Take ten URLs you genuinely care about, run them through each vendor, and read the output. Check the end of long articles, check tables survived, and check that a page that failed is reported as failed rather than returned as an empty success.

Why does per-page pricing vary so much between vendors?

Many price the mechanism, charging multipliers for JavaScript rendering, stealth or premium proxies. Since you cannot know in advance which pages need those, your bill becomes unpredictable. Flat per-page pricing moves that uncertainty to the vendor.

What is the most commonly missed failure mode?

A 200 response containing a challenge or consent page instead of content. It looks like success, embeds cleanly, and quietly poisons your index. Ask every vendor how you would detect it.

Should I just build it myself?

For a handful of known sites with stable markup, often yes. The cost is not the first version, it is keeping it working across thousands of sites that change without telling you.

Try it on your own URLs.

Sign up with a work email for 500 free credits, no card required.

Get API key