All posts

How to monitor a website for changes

September 16th 2026 · Akash Rajpurohit

Website monitoring looks trivial for about a day. Fetch a page, hash it, compare it to yesterday. Then it fires forty times before lunch because a build hash changed, and you learn that detecting a change and detecting a change you care about are different problems.

TLDR

  • Comparing raw HTML produces near-constant false positives. Compare extracted content instead.
  • The expensive question is not “did it change” but “did it change meaningfully”, and that is a reduction problem.
  • Conditional requests answer “did anything change at all” for almost nothing.
  • Match cadence to how fast the page actually moves, not to how often you would like to know.
  • Store what you compared, not just a hash, or you can never explain what changed.

Why hashing the HTML does not work

The obvious implementation is to hash the response and compare. On real pages this fires almost every time, because a modern page contains plenty that differs per request and means nothing:

  • CSRF tokens and session identifiers
  • Build hashes in asset filenames, which change on every deploy
  • Rotating advertisements and recommendation widgets
  • Rendered timestamps, view counters, “3 people are looking at this”
  • A/B test buckets that vary by visitor

None of that is the page changing in any sense you care about. A monitor that alerts on it is worse than no monitor, because people stop reading its alerts within a week and then miss the real one.

Compare the content, not the markup

The fix is to reduce the page to its meaning before comparing. Extract the main content, drop navigation and boilerplate, and diff that.

This removes most noise in one step, because nearly everything on the volatile list above lives in the chrome rather than the article. What remains is text, and text diffs are both stable and readable, which matters when you have to explain to someone what actually changed.

Keep structure while you do it. Headings and tables let you say “the price table changed” rather than “something on this page changed”, and that difference determines whether an alert is actionable.

Narrow further to the part you care about

Content-level diffing still flags things you may not want. A docs page gains a new example; a pricing page adds a footnote. Whether those matter depends entirely on why you are watching.

So define the target more tightly than “the page”:

  • Watching a pricing page? Compare the table, not the copy around it.
  • Watching a status page? Compare current state, not the incident history that grows forever.
  • Watching a competitor’s docs? Compare the section headings to catch new capabilities, rather than every wording tweak.
  • Watching for availability? Compare a specific field, not the document.

The narrower the target, the higher the signal, and monitoring lives or dies on signal because the failure mode is people ignoring it.

Use conditional requests before doing any of this

Before extraction and diffing, there is a much cheaper question: did anything change at all?

Store the ETag and Last-Modified from each fetch, and send them back as If-None-Match and If-Modified-Since. A well-configured server answers 304 Not Modified with no body. No download, no parsing, no diffing, and almost no load on either side.

Not every server implements this properly, so you still need the full path as a fallback. But when it works, the majority of checks on a page that has not changed cost close to nothing, and that is what makes frequent checking affordable at all.

Match the cadence to the page

The instinct is to check often so nothing is missed. In practice most pages change far less than people assume, and over-checking is both wasteful and a good way to get blocked.

A reasonable starting point:

page type cadence
status, availability, stock minutes to hourly
pricing, plans hourly to daily
documentation, changelogs daily
marketing pages, about weekly
archives, old posts monthly, or never

Then adjust from what you observe. If a page has not changed in three months of daily checks, it does not need daily checks. If one changed twice between checks, tighten it. Sites also publish this information themselves: a lastmod in the sitemap and RSS feeds are both cheaper change signals than polling, and both are widely ignored.

Store what you compared

The mistake that hurts later: storing only a hash. When an alert fires, you know something changed and you cannot say what, which is precisely the moment you needed to.

Store the extracted content you compared against. Then a diff is available on demand, alerts can carry the actual change, and you can answer “when did this start” without a rebuild. Extracted text is small compared to raw pages, so this is cheaper than it sounds.

What good monitoring looks like

  1. Conditional request. If 304, stop, and record that you checked.
  2. Otherwise fetch and extract the content.
  3. Narrow to the section that matters.
  4. Compare to the stored version.
  5. If it differs meaningfully, alert with the diff, and store the new version.

Most checks stop at step one. That is what makes the rest affordable.

The part that is genuinely hard

Everything above is scheduling and diffing, which is ordinary engineering. The part that decides whether it works is step two: getting the same page reduced to the same content every time.

An extractor that includes a bit of navigation on some fetches and not others will produce diffs that look real and are not. A pipeline that silently returns a consent wall on one check will report an enormous change and then another enormous change when the real page comes back. Monitoring is only as stable as the extraction underneath it, which is why “compare the content” is easy to say and the actual work.

If you would rather not own that part, scrape returns the same clean markdown for the same page every time, with a quality signal so a walled or empty capture can be discarded instead of diffed. One credit a page, failures never billed, and 500 free credits on a work email to test it against whatever you are watching.

[ FAQ ]

How do I monitor a website for changes?

Fetch the page on a schedule, reduce it to the content you care about, and compare that against what you stored last time. The hard part is the reduction step, because comparing raw HTML flags a change on nearly every check.

Why does my monitor fire constantly when nothing changed?

Because raw HTML contains things that differ on every request, such as CSRF tokens, timestamps, session identifiers, rotating ad slots and build hashes. Compare extracted content rather than markup and most of that noise disappears.

How often should I check a page?

As often as the page actually changes, which is usually far less often than people assume. Pricing and status pages might warrant hourly; a documentation page rarely needs more than daily.

What is the cheapest way to check whether a page changed?

A conditional request. Send If-None-Match or If-Modified-Since and a well-configured server replies 304 with no body, so you skip the download and the parsing entirely.

Should I diff the text or the HTML?

Text, almost always. HTML diffs are dominated by markup churn that has no bearing on what the page says. Diff the extracted content, and keep structure like headings and tables so you can tell where the change happened.

Try it on your own URLs.

Sign up with a work email for 500 free credits, no card required.

Get API key