All posts

How to find every page on a website

September 20th 2026 · Akash Rajpurohit

This sounds like it should be a solved problem, and it is not. Every method of enumerating a site’s pages misses a different category, and the ones that look most authoritative are often the least complete.

TLDR

  • Sitemaps are the fastest and highest-precision start, and they are incomplete by design.
  • Link crawling finds what the site links to, which by definition excludes orphan pages.
  • The site: operator shows an approximate sample of what a search engine indexed, not what exists.
  • Archives find URLs that used to be linked, which is the only practical route to some orphans.
  • No single method is complete. Combine two or three and merge, then decide what you actually need to fetch.

Start with the sitemap, but do not trust it

/sitemap.xml is where a site publishes the URLs it considers canonical. It is the highest-signal source available and it is cheap: one request instead of thousands.

Read robots.txt first, because that is where the sitemap location is usually declared, and large sites often have several sitemaps behind an index file. Sitemaps also frequently carry lastmod, which tells you when each page changed and lets you skip everything that has not.

What sitemaps routinely miss:

  • Pages the site does not want indexed, which are still pages
  • Anything paginated or filtered, since faceted URLs are usually excluded on purpose
  • Recently published pages, if the sitemap regenerates on a schedule
  • Everything, when the sitemap is simply stale, which is common

A sitemap tells you what a site is proud of. That is not the same as what it has.

Following links from the homepage finds what the sitemap missed, including sections that were never registered and pages created after the last regeneration.

The failure mode here is different, and it is structural: link crawling can only ever find pages that something links to. Orphan pages are invisible to it, permanently, no matter how deep you go.

Two practical cautions. Crawl politely, one host at a time with backoff, because enumerating a site is exactly the workload that looks like an attack from the other side. And bound it: an unbounded crawl on a site with faceted navigation can generate effectively infinite URLs from sort and filter parameters, and you will discover this after several hours.

Use the search operator as a cross-check

site:example.com in Google or Bing shows what that engine has indexed for the domain.

Treat this as a comparison, never a source. The result counts are estimates and famously imprecise, you cannot page through them exhaustively, and anything the engine chose not to index simply is not there.

Where it genuinely earns its place is as a diff. If the operator surfaces a section your crawl never reached, you have found a gap in your discovery, and that is worth knowing even though the operator itself is a poor enumerator.

Archives find the orphans

The Wayback Machine and similar archives hold URLs that existed at some point, which makes them the one practical way to find pages nothing links to any more.

This matters for a specific and common job: auditing what happened to a site after a redesign. Pages that were dropped, moved without redirects, or quietly unpublished show up in an archive and in no live source at all.

The tradeoff is that archives are historical. Plenty of what you find will be dead, so anything you pull from here needs verifying before you act on it.

What each method actually finds

method finds misses
Sitemap canonical, indexable pages orphans, excluded, stale entries
Link crawl anything reachable by link orphan pages, entirely
site: operator an indexed sample unindexed pages, exact counts
Archives historical URLs anything new; includes dead URLs

The overlap is large and the gaps do not coincide, which is why merging two or three gets you most of the way and any single one leaves a hole.

Map first, then decide what to fetch

The mistake worth avoiding is conflating two jobs. Discovery is finding out which URLs exist. Retrieval is fetching their content. They have completely different costs, and doing them together means paying retrieval prices to answer a discovery question.

Map first. Look at the shape: how many pages, which sections, what the URL patterns are. Then choose the subset worth fetching. On most sites you will find you want a fraction of what exists, and you will have learned that for the price of a handful of requests rather than thousands.

This is also the honest answer to “how many pages does this site have”. The number depends on which methods you combined, and anyone quoting a single confident figure has picked one method and not told you which.

Doing it in one call

If you would rather not assemble this yourself, map does the discovery half: give it a domain and it returns the URLs, sitemap-aware, for 1 credit regardless of how many it finds. Then crawl fetches only the parts you chose, at one credit a page, with failures never billed.

A work email gets you 500 free credits, which is enough to map a few sites and see their real shape before you commit to fetching anything.

[ FAQ ]

How do I find all the pages on a website?

Start with the sitemap at /sitemap.xml, which lists the URLs the site considers canonical. Then crawl the links to catch pages the sitemap omits. No single method finds everything, so the practical answer is to combine two or three and merge the results.

Does a sitemap list every page?

No. A sitemap lists the pages a site wants indexed. Pages that are unlinked, paginated, filtered, recently added or deliberately excluded are routinely missing, and plenty of sitemaps are stale.

Is the Google site: operator reliable for this?

It is useful as a cross-check and unreliable as a source of truth. It shows a sample of what Google has indexed, not what exists, the counts are approximate, and it silently omits anything Google chose not to index.

How do I find pages that are not linked from anywhere?

Orphan pages will not appear from link crawling by definition. The sitemap sometimes has them, and archives like the Wayback Machine often hold URLs that were linked once and are not any more.

What is the difference between mapping and crawling a site?

Mapping returns the list of URLs without fetching each page, which is fast and cheap. Crawling fetches the content of each page. Map first to see the shape and size of a site, then crawl the parts you actually want.

Try it on your own URLs.

Sign up with a work email for 500 free credits, no card required.

Get API key