Your index is only as good as what you put in it.
Retrieval fails on the ingestion side long before it fails on the model side. Map a site to every URL it has, pull each page as clean markdown, and chunk on structure the page already carries.
one site · one pass
[ 01 / How it works ]
Three calls. One index.
The order matters. Mapping first turns ingestion into a list you can diff on the next run, instead of a crawl you restart and hope covered everything.
Map
Find every page first
One call returns the whole URL set for a site, so ingestion works from a list you can diff rather than a crawl you hope finished.
const { links } = await hf.map({
url: site,
});Crawl
Pull clean markdown
Navigation, cookie banners and footers are gone before you see the page. Headings, tables and code blocks survive, because chunking depends on them.
const pages = await hf.crawl({
url: site,
limit: 500,
});Embed
Chunk on the headings
Split at the structure the page already has, keep the heading path on every chunk, and a retrieved fragment still knows where it came from.
for (const page of pages) {
await index.upsert(
chunk(page.markdown),
);
}[ 02 / The demo ]
One site, traced from URL to vector.
A real run, shown as one pass rather than three screens. The site on the left narrows to the pages worth reading, and one of those is followed all the way down to the records that get embedded.
01Map1 credit
02Crawl4 credits
03Chunkfree, in your code
fastapi.tiangolo.com
├── async
├── features
├── learn
├── python-types
└── tutorial/
├── body
├── body-fields
├── body-multiple-params
├── body-nested-models
├── body-updates
├── cookie-params
├── dependencies/
│ └── sub-dependencies
├── encoder
├── extra-data-types
├── extra-models
├── first-steps
├── handling-errors
├── header-params
├── path-operation-configuration
├── path-params
├── path-params-numeric-validations
├── query-param-models
├── query-params
├── query-params-str-validations
├── request-files
├── request-forms
├── response-model
├── response-status-code
├── schema-extra-example
└── security
+119 more
first-steps1,767w
path-params1,474w
handling-errors1,662w
query-params836w
followed below
markdown returned5,739w
nav, sidebar, footer and cookie banner dropped · headings, code and tables kept
Query Parameters163w
Query Parameters
When you declare other function parameters that are not part of the path parameters, they are automatically interpreted as "query" parameters. ``` fr
Defaults89w
Query Parameters › Defaults
As query parameters are not a fixed part of a path, they can be optional and can have default values. In the example above they have default values
Optional parameters92w
Query Parameters › Optional parameters
The same way, you can declare optional query parameters, by setting their default to `None`: ``` from fastapi import FastAPI app = FastAPI() @app.ge
Query parameter type conversion116w
Query Parameters › Query parameter type conversion
You can also declare `bool` types, and they will be converted: ``` from fastapi import FastAPI app = FastAPI() @app.get("/items/{item_id}") async de
Multiple path and query parameters97w
Query Parameters › Multiple path and query parameters
You can declare multiple path parameters and query parameters at the same time, **FastAPI** knows which is which. And you don't have to declare them
Required query parameters276w
Query Parameters › Required query parameters
When you declare a default value for non-path parameters (for now, we have only seen query parameters), then it is not required. If you don't want t
/tutorial/query-params · 836 words · 6 chunks · every one carries its heading path
The heading path is the point. Every record on the right carries the page title and the section it came from, so a fragment retrieved on its own still says what it is about. Split the same page every eight hundred characters instead and the fifth record is a block of Python with nothing left in it to say which framework it belongs to.
[ 03 / Built for ]
Anyone answering questions from someone else's pages.
The retrieval problem is the same whether the corpus is your own docs, a regulator's website, or a few hundred competitors.
Docs assistants
The assistant answers from a snapshot taken before the last release.
Re-map on a schedule and only the pages that moved get fetched again.
Developer tools
Code samples arrive mangled, so the model suggests syntax that will not run.
Code blocks come back fenced and intact, language and indentation included.
Regulated industries
An answer without a source is an answer nobody is allowed to act on.
Every chunk keeps its URL and heading path, so a citation is already there.
Knowledge bases
Content lives across a wiki, a help centre and a marketing site.
One call per site, the same markdown out, one pipeline to maintain.
Research agents
An agent handed raw HTML spends its context window on markup.
Clean text means the window holds the corpus, not the page furniture.
Vertical search
Indexing a whole sector means writing a parser per site.
The same call handles every site, so coverage stops being an engineering cost.
[ 04 / Keep going ]
Same API. Other problems.
Web data for AI
Give an agent the live web
Search mid-answer and get results already fetched, cleaned and citable.
ReadWeb data for AI
Turn listings into a dataset
Point a JSON schema at a directory and get typed rows back.
ReadWeb data for AI
Watch pages for changes
Re-run a set on a schedule and diff what came back.
Read[ Start ]
Clean web data is one call away.
500 free credits, no card required. Failures are never billed.
Success rate
Median scrape
ms