llms.txt vs robots.txt
Access control versus reading guidance — different files, different authority.
Short answer: robots.txt answers “may you fetch this?” — it is an access-control convention for crawlers, with real semantics and an RFC behind it. llms.txt answers “what is worth reading?” — a recommendation with no force at all. They are not substitutes and never were: one restricts, the other curates. Keeping both is the normal setup for a site that wants AI visibility it can also bound.
What robots.txt does
RFC 9309 formalized the Robots Exclusion Protocol that crawlers have honored since the 1990s. The file lives at/robots.txt and groups Disallow and Allow path rules under User-agent tokens:
User-agent: *
Disallow: /admin/
Disallow: /cart
Sitemap: https://acme.com/sitemap.xmlCompliance is voluntary — a well-behaved crawler obeys, a scraper ignores — but the major search and AI crawlers document their support, and Google’s crawler documentationdescribes exactly how rules are parsed and grouped. It is the closest thing the web has to an enforced front door.
What llms.txt does
llms.txt is a markdown suggestion, not a directive. It lists the pages an AI should read — with descriptions — under readable headings. No crawler is obliged to fetch it, obey it, or even know it exists. Its value is entirely in what it tells a cooperative reader: the agent that lands on your site wanting to do a good job gets your own summary of what matters, in one request, instead of crawling blind.
Side-by-side comparison
| robots.txt | llms.txt | |
|---|---|---|
| Purpose | Restrict crawling | Guide reading |
| Grammar | User-agent / Disallow / Allow rules (RFC 9309) | Markdown: H1, blockquote, link sections |
| Force | Documented, near-universal crawler support | None — pure suggestion |
| Granularity | Path patterns, allow/deny | Curated page list with descriptions |
| Who reads it | Crawlers, before fetching | AI agents, deciding what to fetch |
| Standards status | IETF RFC (2022) | Open proposal (llmstxt.org) |
The most common llms.txt misconception
The recurring mistake is treating llms.txt as a robot ban: a site owner publishes one, AI traffic continues anyway, and the file gets blamed for “not working”. llms.txt cannot exclude anyone — it is a list of things to read, not a fence. The direction of the complaint gives it away: robots.txt problems are about too much access; llms.txt problems are about too little understanding. If your goal is to keep bots out, edit robots.txt (and check each provider’s crawler docs for their user-agent token); if your goal is to be understood, write a good llms.txt.
Do you need both?
| Your goal | robots.txt | llms.txt |
|---|---|---|
| Welcome AI, help it read well | Keep permissive | Curate carefully |
| Keep some areas private | Disallow those paths | Simply don’t list them |
| Opt out of AI entirely | Disallow major AI crawlers | Skip the file |
| Ship a sitemap for search | Keep the Sitemap: line | Unrelated — keep it anyway |
What robots.txt does better — honestly
robots.txt has thirty years of tooling, an RFC, and documented compliance from the crawlers that matter. llms.txt has a young proposal and voluntary adoption. If you take one lesson from this page: use robots.txt for anything with consequences (access, privacy, load), and treat llms.txt as presentation — the front desk, not the lock. When you are ready to write one, the llms.txt generator builds it from your live site and shows every inclusion decision. Related reading: llms.txt vs Schema.orgfor the structured-data layer.
Common mistakes
- Putting Disallow-style rules in llms.txt or expecting it to reduce traffic.
- Blocking AI crawlers in robots.txt while publishing a detailed llms.txt — the second is unread the moment the first applies.
- Forgetting the Sitemap: directive when adding either file; it still serves search crawling.