Your index is only as good as what you put in it.

Retrieval fails on the ingestion side long before it fails on the model side. Map a site to every URL it has, pull each page as clean markdown, and chunk on structure the page already carries.

Get API key
fastapi.tiangolo.com200
urls mapped151
pages crawled4
words returned5,739
markup strippednav · footer · banners
structure keptheadings · code · tables
chunks from one page6

one site · one pass

[ 01 / How it works ]

Three calls. One index.

The order matters. Mapping first turns ingestion into a list you can diff on the next run, instead of a crawl you restart and hope covered everything.

Map

Find every page first

One call returns the whole URL set for a site, so ingestion works from a list you can diff rather than a crawl you hope finished.

const { links } = await hf.map({
  url: site,
});

Crawl

Pull clean markdown

Navigation, cookie banners and footers are gone before you see the page. Headings, tables and code blocks survive, because chunking depends on them.

const pages = await hf.crawl({
  url: site,
  limit: 500,
});

Embed

Chunk on the headings

Split at the structure the page already has, keep the heading path on every chunk, and a retrieved fragment still knows where it came from.

for (const page of pages) {
  await index.upsert(
    chunk(page.markdown),
  );
}

[ 02 / The demo ]

One site, traced from URL to vector.

A real run, shown as one pass rather than three screens. The site on the left narrows to the pages worth reading, and one of those is followed all the way down to the records that get embedded.

fastapi.tiangolo.com151 urls→4 pages→6 records

01Map1 credit

02Crawl4 credits

03Chunkfree, in your code

fastapi.tiangolo.com

├── async

├── features

├── learn

├── python-types

└── tutorial/

├── body

├── body-fields

├── body-multiple-params

├── body-nested-models

├── body-updates

├── cookie-params

├── dependencies/

│ └── sub-dependencies

├── encoder

├── extra-data-types

├── extra-models

├── first-steps

├── handling-errors

├── header-params

├── path-operation-configuration

├── path-params

├── path-params-numeric-validations

├── query-param-models

├── query-params

├── query-params-str-validations

├── request-files

├── request-forms

├── response-model

├── response-status-code

├── schema-extra-example

└── security

+119 more

first-steps1,767w

path-params1,474w

handling-errors1,662w

query-params836w

followed below

markdown returned5,739w

nav, sidebar, footer and cookie banner dropped · headings, code and tables kept

Query Parameters163w

Query Parameters

When you declare other function parameters that are not part of the path parameters, they are automatically interpreted as "query" parameters. ``` fr

Defaults89w

Query Parameters › Defaults

As query parameters are not a fixed part of a path, they can be optional and can have default values. In the example above they have default values

Optional parameters92w

Query Parameters › Optional parameters

The same way, you can declare optional query parameters, by setting their default to `None`: ``` from fastapi import FastAPI app = FastAPI() @app.ge

Query parameter type conversion116w

Query Parameters › Query parameter type conversion

You can also declare `bool` types, and they will be converted: ``` from fastapi import FastAPI app = FastAPI() @app.get("/items/{item_id}") async de

Multiple path and query parameters97w

Query Parameters › Multiple path and query parameters

You can declare multiple path parameters and query parameters at the same time, **FastAPI** knows which is which. And you don't have to declare them

Required query parameters276w

Query Parameters › Required query parameters

When you declare a default value for non-path parameters (for now, we have only seen query parameters), then it is not required. If you don't want t

/tutorial/query-params · 836 words · 6 chunks · every one carries its heading path

The heading path is the point. Every record on the right carries the page title and the section it came from, so a fragment retrieved on its own still says what it is about. Split the same page every eight hundred characters instead and the fifth record is a block of Python with nothing left in it to say which framework it belongs to.

[ 03 / Built for ]

Anyone answering questions from someone else's pages.

The retrieval problem is the same whether the corpus is your own docs, a regulator's website, or a few hundred competitors.

Docs assistants

The assistant answers from a snapshot taken before the last release.

Re-map on a schedule and only the pages that moved get fetched again.

Developer tools

Code samples arrive mangled, so the model suggests syntax that will not run.

Code blocks come back fenced and intact, language and indentation included.

Regulated industries

An answer without a source is an answer nobody is allowed to act on.

Every chunk keeps its URL and heading path, so a citation is already there.

Knowledge bases

Content lives across a wiki, a help centre and a marketing site.

One call per site, the same markdown out, one pipeline to maintain.

Research agents

An agent handed raw HTML spends its context window on markup.

Clean text means the window holds the corpus, not the page furniture.

Vertical search

Indexing a whole sector means writing a parser per site.

The same call handles every site, so coverage stops being an engineering cost.

[ 04 / Keep going ]

Same API. Other problems.

[ Start ]

Clean web data is one call away.

500 free credits, no card required. Failures are never billed.

Success rate

 

Median scrape

 ms