All posts

How to make your site readable by AI agents

September 7th 2026 · Akash Rajpurohit

There is a lot of noise right now about making sites legible to AI. Some of it is real engineering, some is the same people who sold meta keyword stuffing wearing a new hat. This post separates the two, and is written from having actually implemented most of it.

TLDR

  • robots.txt is the only item here that is a genuine standard with broad compliance. Start there, and decide deliberately which AI crawlers you allow.
  • llms.txt is a useful convention, not a standard. Adding it costs an hour and it will not hurt you; expecting crawlers to obey it will.
  • agents.md is for agents that intend to act, not just read. If you have an API, this is the highest-value file on the list.
  • Serving markdown to clients that ask for it is the most concretely useful thing you can do, because it removes a lossy step from the consumer’s pipeline.
  • None of this is a ranking trick. It changes whether a model that already found you can use you correctly.

Start with the thing that is actually a standard

robots.txt is decades old, universally understood and broadly honoured by the crawlers that matter. Every other file in this post is advisory. This one is the real control surface, and the interesting part now is that you have to make a decision you did not used to have to make: which AI crawlers do you want?

The bots split into rough groups. Some fetch pages to train models. Some fetch pages live, to answer a question a user is asking right now, and cite the source. Some do both. Blanket-blocking everything with “AI” in the name will also block the ones that send people to you.

Whatever you decide, decide it explicitly rather than by inheriting a template. And be aware that a Disallow is a request, not a wall: it stops well-behaved crawlers and does nothing to anyone else.

What llms.txt is, and what it is not

llms.txt is a markdown file at your site root that gives a model a curated tour: what this site is, what the main sections are, which pages actually matter.

The honest framing is that it is a convention with real momentum but no obligation attached. No crawler is required to fetch it. No search engine has committed to weighting it. It is a bet that as models increasingly navigate sites directly, a short human-curated map beats making them infer your information architecture from a nav bar.

That bet looks reasonable, and the cost is an hour of work, so we wrote one. What we did not do is expect it to change anything measurable on its own. If someone is selling you llms.txt as a traffic strategy, that is the tell.

The practical advice: keep it short, link to the pages you would actually want quoted, and describe things in plain language rather than marketing copy. A model reading “the leading AI-powered platform for synergistic data solutions” learns nothing it can use in an answer.

agents.md is the one worth your time

If llms.txt orients a reader, agents.md briefs something that intends to do something. It answers the questions an agent has to resolve before it can act on your behalf:

  • What does this product actually do, in operational terms?
  • Where is the API, and what shape are the requests?
  • How does authentication work, and how does an agent obtain a credential?
  • What are the rate limits and what happens when I hit them?
  • What does an error look like?

If you have an API, this is the highest-leverage file on the list, because the alternative is an agent guessing from your marketing pages and getting it wrong. It also has a pleasant side effect: writing it honestly exposes every place your API is harder to use than you thought.

Pair it with a machine-readable API description. An OpenAPI document an agent can fetch without authenticating is worth more than any amount of prose, because it removes the guessing entirely.

Serving markdown to clients that ask for it

This is the most concretely useful item here, and the least discussed.

When an agent fetches your page, it gets HTML, and then it has to extract the content from that HTML. That step is lossy and it is not free. If your server notices a client asking for text/markdown and hands back clean markdown directly, you have removed an entire error-prone stage from their pipeline, and you control the output rather than hoping their extractor guesses right about your templates.

The same applies to letting a human append something like .md to a URL. It costs little and it means anyone quoting your documentation is quoting the version you wrote, not a mangled one.

What about signalling how your content may be used?

There are competing proposals for expressing usage preferences in a machine-readable way, some through robots.txt directives, some through response headers. They are worth adding because they are cheap and they express intent clearly.

Be realistic about enforcement. These are declarations, not access controls. They give well-behaved consumers a clear signal to respect and give you a documented position, which matters more than it might appear if the question is ever contested. What they do not do is prevent anybody from doing anything.

A sensible order to do this in

  1. robots.txt, with a deliberate decision about AI crawlers rather than a copied default.
  2. A sitemap, submitted. Unglamorous and still the main way anything gets discovered.
  3. agents.md plus a fetchable OpenAPI document, if you have an API.
  4. Markdown responses for clients that ask for them.
  5. llms.txt, written like a map rather than a brochure.
  6. Usage signals, as a clear statement of intent.

Roughly a day’s work in total, and the first three carry most of the value.

The part nobody wants to hear

None of this makes a model discover you. It makes a model that has already reached you able to read you accurately, quote you correctly and act against your API without guessing.

Discovery is still the hard problem, and it is still mostly solved by being genuinely useful somewhere a model already reads: real documentation, real repositories, real conversations. Files at your site root are hygiene. They are worth doing because they are cheap and correct, not because they are a growth channel.

Most of the checklist above is quicker to verify than to reason about. The AI readability audit reads a few of your pages the way an agent does and reports what it finds: whether your content survives without JavaScript, whether robots.txt lets assistants in, whether you publish an llms.txt, and how much of what you serve is markup rather than text. It needs no signup, and it shows you the markdown an agent actually receives for your page, which is usually the point at which the problem stops being abstract. If the missing piece turns out to be llms.txt, the llms.txt generator drafts one from your own sitemap.

If you are on the other side of this problem and need to consume the web rather than be consumed, that is what we build: the scrape endpoint turns any URL into clean markdown at one credit a page, and a work email gets you 500 free credits. You can also try it with no signup in the URL to markdown tool.

[ FAQ ]

What is llms.txt?

A markdown file at the root of your site that gives a language model a curated map of what you offer and where the important pages are. It is a community convention rather than a ratified standard, and no crawler is obliged to read it.

Is llms.txt the same as robots.txt?

No. robots.txt tells crawlers what they may fetch and is widely honoured. llms.txt tells a model what is worth reading and is advisory. One is about permission, the other about orientation.

What is agents.md?

A file describing how an agent should work with your product: what the API does, how to authenticate, what the limits are. Where llms.txt orients a reader, agents.md is written for something that intends to act.

Do I need to serve markdown versions of my pages?

It helps. An agent asking for text/markdown and getting clean markdown skips the extraction step entirely, which makes your content cheaper to consume and less likely to be mangled.

Does any of this improve my ranking?

Not directly, and be sceptical of anyone who promises it does. What it changes is whether a model that reaches your site can understand and cite it accurately.

Try it on your own URLs.

Sign up with a work email for 500 free credits, no card required.

Get API key