All posts

What web scraping actually costs

August 27th 2026 · Akash Rajpurohit

Ask what web scraping costs and you get a per-page price. That number is real, and it is usually the smallest line in the budget. The costs that decide whether a project is viable are the ones nobody quotes.

TLDR

  • The per-page fee is the visible cost and rarely the dominant one.
  • Maintenance is the real expense when you build, and it does not decline over time.
  • Multiplier pricing makes a vendor bill unpredictable, which is worse than it being higher.
  • Failure handling is a cost centre: retries, billed failures, and bad data reaching production.
  • Estimate on your own URLs. Category averages predict nothing about your particular mix.

The costs of building it yourself

The first version is cheap and fast, which is exactly what makes this decision easy to get wrong. A working extractor for a known site is an afternoon. The costs arrive later.

Maintenance is the big one. Sites redesign without telling you. A selector that worked for a year breaks silently, and the failure mode is not an error, it is an empty field that flows into your database looking like an absence rather than a fault. Across a handful of sites this is a manageable annoyance. Across hundreds it is somebody’s recurring job, forever.

Infrastructure. Fetching at volume means somewhere to run it, and pages that only exist after JavaScript runs need considerably more than an HTTP client. Both cost money and, more importantly, attention.

The long tail of awkward sites. A predictable share of any real target list will resist straightforward fetching. Each one is an investigation, and investigations are the least schedulable work there is.

Storage. Raw captures accumulate. Keeping them is genuinely valuable, because it lets you re-extract without refetching when your parsing improves, but it is a line item that only grows.

None of this argues against building. It argues for counting it. A build decision made on the cost of the first version is a decision made on the smallest number in the problem.

The costs of buying

The per-page fee, which is the number on the pricing page and the one people compare.

Multipliers, which are the number that actually determines your bill. Many vendors price the mechanism: a base rate for a simple fetch, then more for rendering, more for stealth, more for premium proxies, sometimes compounding.

The issue is not that this is expensive, it is that it is unpredictable. You cannot know in advance which of your URLs will need which treatment, because that is a property of each site and it changes when the site changes. Your costs then move for reasons you cannot see or control, which makes forecasting impossible and makes a cheap-looking headline rate meaningless.

Flat per-page pricing moves that variance to the vendor. We price this way deliberately: a page costs one credit whatever it took to fetch. Additional product costs more, so model-backed extraction is priced above a plain fetch because you chose it and it carries real marginal cost. But how hard the page was to get is not your problem.

Being billed for failures. Ask directly. Paying for a request that returned a verification page means paying for someone else’s anti-bot vendor, and at volume it is a real number. Our position is that failures cost nothing.

The cost everyone forgets

Bad data is more expensive than no data, and it does not appear in any pricing comparison.

A failed fetch that returns an error is cheap: you retry it, or you log it and move on. A fetch that returns a 200 containing a consent wall is expensive, because nothing catches it. It embeds cleanly, indexes cleanly, and sits there until it produces a confident wrong answer in front of a user or a decision.

The cost of that is not the credit. It is the hours spent tracing back why an answer was wrong, plus whatever the wrong answer caused. This is why a quality signal on every response is a cost feature and not a nicety: it is the difference between catching that at ingestion and catching it in production.

Working out your own number

Category averages are useless here, because cost depends almost entirely on your particular mix of sites. Do this instead:

  1. Take a representative sample of your real URLs, including the awkward ones. Fifty is plenty.
  2. Run them through each option you are considering.
  3. Count what you would actually be charged, including failures and retries, not the headline rate.
  4. Add a maintenance estimate if building. Be honest: how many hours a month, at what rate, and remember it grows with the number of sites.
  5. Multiply by your real volume, not your ambition.

That gives you two numbers you can defend. Most of the argument between building and buying disappears once both are written down, because they are usually not close.

Where the line generally falls

Build when you have a small number of stable, well-behaved sites, when the extraction is genuinely core to your product, or when you need control that no API gives you.

Buy when the site list is long or changing, when extraction is not the thing you are actually selling, or when the cost of a quiet failure exceeds the cost of the fetch.

The honest summary: the per-page price is the part you can see, and rarely the part that decides. Cost is dominated by what happens when a page does not come back cleanly, and by whose job it is to fix that.

If you want a number for your own workload, the scrape endpoint is one credit a page whatever it took to fetch, with failures never billed and a quality signal on every response. A work email gets you 500 free credits, which is enough to run a real sample and get a figure you can actually plan against.

[ FAQ ]

How much does web scraping cost?

The per-page fee is usually the smallest part. The real cost is engineering time to build and then maintain extractors as sites change, plus the compute and infrastructure behind fetching pages that resist being fetched.

Is it cheaper to build my own scraper?

For a small number of stable sites, usually yes. Across many sites that change without warning, rarely, because the cost is maintenance rather than the first version and maintenance does not stop.

Why do web scraping bills vary so much between vendors?

Because many price the mechanism, charging multipliers for rendering, stealth or premium proxies. Since you cannot predict which pages need those, the effective rate is unknowable in advance.

What hidden costs do people miss?

Failed requests you are billed for, retries, storage of raw captures, the engineering hours spent diagnosing why a site broke, and the cost of bad data reaching production and being acted on.

How do I estimate my real cost before committing?

Take a representative sample of your actual URLs, run them through each option, and count what you would be charged including failures. Then add a realistic maintenance figure, because that is the line that grows.

Try it on your own URLs.

Sign up with a work email for 500 free credits, no card required.

Get API key