llms.txt vs robots.txt

Access control versus reading guidance — different files, different authority.

Short answer: robots.txt answers “may you fetch this?” — it is an access-control convention for crawlers, with real semantics and an RFC behind it. llms.txt answers “what is worth reading?” — a recommendation with no force at all. They are not substitutes and never were: one restricts, the other curates. Keeping both is the normal setup for a site that wants AI visibility it can also bound.

What robots.txt does

RFC 9309 formalized the Robots Exclusion Protocol that crawlers have honored since the 1990s. The file lives at/robots.txt and groups Disallow and Allow path rules under User-agent tokens:

User-agent: *
Disallow: /admin/
Disallow: /cart

Sitemap: https://acme.com/sitemap.xml

Compliance is voluntary — a well-behaved crawler obeys, a scraper ignores — but the major search and AI crawlers document their support, and Google’s crawler documentationdescribes exactly how rules are parsed and grouped. It is the closest thing the web has to an enforced front door.

What llms.txt does

llms.txt is a markdown suggestion, not a directive. It lists the pages an AI should read — with descriptions — under readable headings. No crawler is obliged to fetch it, obey it, or even know it exists. Its value is entirely in what it tells a cooperative reader: the agent that lands on your site wanting to do a good job gets your own summary of what matters, in one request, instead of crawling blind.

Side-by-side comparison

robots.txtllms.txt
PurposeRestrict crawlingGuide reading
GrammarUser-agent / Disallow / Allow rules (RFC 9309)Markdown: H1, blockquote, link sections
ForceDocumented, near-universal crawler supportNone — pure suggestion
GranularityPath patterns, allow/denyCurated page list with descriptions
Who reads itCrawlers, before fetchingAI agents, deciding what to fetch
Standards statusIETF RFC (2022)Open proposal (llmstxt.org)

The most common llms.txt misconception

The recurring mistake is treating llms.txt as a robot ban: a site owner publishes one, AI traffic continues anyway, and the file gets blamed for “not working”. llms.txt cannot exclude anyone — it is a list of things to read, not a fence. The direction of the complaint gives it away: robots.txt problems are about too much access; llms.txt problems are about too little understanding. If your goal is to keep bots out, edit robots.txt (and check each provider’s crawler docs for their user-agent token); if your goal is to be understood, write a good llms.txt.

Do you need both?

Your goalrobots.txtllms.txt
Welcome AI, help it read wellKeep permissiveCurate carefully
Keep some areas privateDisallow those pathsSimply don’t list them
Opt out of AI entirelyDisallow major AI crawlersSkip the file
Ship a sitemap for searchKeep the Sitemap: lineUnrelated — keep it anyway

What robots.txt does better — honestly

robots.txt has thirty years of tooling, an RFC, and documented compliance from the crawlers that matter. llms.txt has a young proposal and voluntary adoption. If you take one lesson from this page: use robots.txt for anything with consequences (access, privacy, load), and treat llms.txt as presentation — the front desk, not the lock. When you are ready to write one, the llms.txt generator builds it from your live site and shows every inclusion decision. Related reading: llms.txt vs Schema.orgfor the structured-data layer.

Common mistakes

FAQ

Do I need llms.txt if I already have robots.txt?
They do different jobs, so most sites end up with both. robots.txt manages access for crawlers; llms.txt curates reading for AI agents. Having one says nothing about the other.
Can llms.txt block AI from my site?
No. llms.txt has no enforcement — it is a suggestion, not a rule. To request that crawlers stay out, use robots.txt Disallow rules, and note that compliance is a per-crawler policy, not a technical guarantee.
Does llms.txt replace the Sitemap directive in robots.txt?
No. The Sitemap: line points crawlers at your URL list and remains important for search. llms.txt is a curated, opinionated subset aimed at AI agents, not a replacement for a sitemap.
Do AI crawlers respect robots.txt?
The major providers document that their crawlers honor robots.txt, with per-bot user-agent tokens you can target. Verify against each provider’s crawler documentation — behavior and token names are provider-specific.
Where does each file live?
Both live at the site root: robots.txt at /robots.txt and llms.txt at /llms.txt. Neither supports relative links usefully — robots.txt paths are root-relative, and llms.txt links should be absolute URLs.