How to build a chatbot over your documentation
September 23rd 2026 · Akash Rajpurohit
Docs chatbots have a reputation for being disappointing, and it is rarely the model’s fault. The retrieval stack is commodity now. What separates a bot people use from one they abandon after a week is almost entirely what happened before anything was embedded.
TLDR
- Ingestion decides quality. Chunking and retrieval are standard; the input rarely is.
- Crawl the rendered site, not the repository, unless you have a specific reason not to.
- Split on headings. Documentation is already structured; ignoring that structure throws away the best chunk boundaries you will ever get.
- Keep code blocks intact. A split code sample is worse than no code sample.
- Re-ingest on deploy. A confidently stale answer is the failure mode users remember.
Crawl the site, not the repository
The instinct is to point at the markdown in your repo, since it is right there and already clean. It is usually the wrong source.
The rendered documentation is what your users actually read. It includes generated API references that do not exist as source files, partials resolved into place, and version selectors applied. Repository markdown is a template, and the gap between template and rendered page is exactly where wrong answers come from.
There is also a maintenance argument. Crawling the site keeps working when the docs move, get restructured, or migrate to a different generator. A repo-path-based ingestion breaks silently on all three.
Start from the sitemap, which documentation generators almost always produce, and you get the page list without crawling blindly.
Split on headings, not on character counts
Fixed-size chunking is the default in most tutorials and it is a poor fit for documentation, because it cuts through the middle of explanations and separates a heading from the thing it introduces.
Documentation is already structured. Every ## is an author telling you where a self-contained idea starts. Split there and each chunk is a coherent unit that begins by saying what it is about, which is exactly what makes retrieval work.
Practical shape:
- Split at
##and### - If a section is very long, split further at paragraph boundaries rather than mid-sentence
- Prefix every chunk with its page title and heading path. “Authentication > Rotating keys” at the top of a chunk gives the retriever an enormous amount to match on, and costs a handful of tokens
- Never split a code block
That last one matters more than it sounds. Half a code sample retrieved on its own is worse than none, because it looks like an answer.
Keep the structure the docs already have
Flattening pages to plain text is the other common ingestion mistake, and it destroys the things documentation depends on:
- Tables of parameters or fields become an unreadable run of words
- Code blocks lose their whitespace and their boundaries
- Lists stop being distinguishable from prose
- Links to related pages disappear, and those are often the correct answer
Markdown preserves all of it and models read it natively. There is no reason to throw it away on the way in.
Strip the navigation, or embed it a thousand times
Every documentation page carries the same sidebar, the same header, the same footer. Ingest the raw page and every chunk contains a slice of that.
Two things go wrong. Chunks get inflated with text that has no bearing on their content, and worse, every page starts to look like every other page, because they genuinely share most of their text. Retrieval quality drops for a reason that never appears in your retrieval code.
Re-ingest on deploy
A docs bot answering from stale documentation is worse than no bot, because the answer looks exactly as authoritative as a correct one and the user has no way to tell.
Trigger ingestion from your docs deploy if you can. If not, run it daily. Use the sitemap’s lastmod and conditional requests so unchanged pages cost nothing, which makes frequent re-ingestion cheap enough to be boring.
Make refusal the default
The final piece is in the prompt, and it is one line most implementations skip:
Answer only from the documentation below. If it does not
cover the question, say so and point to the closest page.
Without it, a model handed three loosely related chunks will produce a confident answer built from general knowledge about similar products. That is the single most damaging thing a docs bot can do, because it is indistinguishable from a real answer until someone follows it.
Pair the instruction with a check on the retrieval itself: if nothing scored above a sensible threshold, do not call the model at all. Say you do not know and offer search. Users forgive “I don’t know” and do not forgive being confidently misled.
The order that actually matters
- Crawl the rendered docs to clean markdown
- Split on headings, prefixed with the heading path
- Embed, retrieve, and answer only from what came back
- Re-ingest on deploy
- Refuse when the context is weak
Steps two through five are a day of work. Step one is where the quality is, and it is the step people spend the least time on.
If you want it handled, crawl turns a documentation site into clean markdown with the navigation and boilerplate already gone, one credit a page and failures never billed. Every response carries a quality signal, so a page that came back empty can be dropped before it reaches your index instead of after. A work email gets you 500 free credits, enough to ingest a real docs site and see what the chunks look like.
[ FAQ ]
How do I build a chatbot over my documentation?
Crawl the docs to clean markdown, split on headings, embed the chunks, retrieve on a question and answer only from what was retrieved. The retrieval machinery is standard; the quality comes almost entirely from the ingestion step.
Should I use the rendered site or the source markdown?
The rendered site, usually. It is what your users actually read, it includes generated API references, and it stays correct when the site changes without anyone updating the repository.
How often should I re-ingest the docs?
On deploy if you can trigger it, daily otherwise. A docs bot answering from last month's documentation is worse than no bot, because the answer looks current.
Why does my bot answer confidently about things that are not in the docs?
Because nothing told it not to. A model handed weak context will still produce fluent text, so the instruction to refuse and the check that context actually arrived both have to be explicit.
What is the most common cause of a bad docs bot?
Ingestion. Navigation sidebars embedded into every chunk, code blocks flattened into prose, and pages that came back empty are all invisible at ingestion time and obvious in every answer afterwards.
Try it on your own URLs.
Sign up with a work email for 500 free credits, no card required.