llms.txt is a plain-text file at the root of your domain that tells a large language model what your site is and which pages are worth reading. It is written in Markdown, lives at yourdomain.com/llms.txt next to robots.txt, and was proposed by Jeremy Howard in September 2024. The format is published at llmstxt.org.
Updated 4 October 2026.
What goes in the file
The specification is deliberately small. A file has four parts, and only the first is required.
- An H1 with your site or company name. One line beginning with
#. This is the only required element in the whole format. - A blockquote summary. A paragraph beginning with
>saying what you are and who you are for. If a model reads one line of your file, it reads this one. - Optional detail paragraphs. Plain prose for what the summary cannot carry: where you are based, how pricing works, how you want to be attributed.
- H2 sections of links. Each groups related pages under a heading such as
## Start here. Every bullet is a Markdown link followed by a colon and a short description. An H2 named## Optionalis reserved by convention for secondary links a model can skip when it needs a shorter context.
The descriptions after each link are the part people skip and the part that does the work. A bare list of URLs tells a model nothing your sitemap does not already say.
The honest answer on whether it does anything
As of October 2026, no major AI platform documents reading llms.txt, and Google’s Search Central documentation says plainly that Google Search ignores them. There is no published evidence that having one changes whether you get cited, and anyone selling it as a ranking factor is ahead of the evidence.
That is the part most explainers leave out, so it is worth being clear about what the file is not. It is not a ranking signal. It is not a permission system, and it carries no legal weight. It does not stop anyone training on your content, which is a robots.txt job. And publishing one grants nothing: if robots.txt blocks a crawler, listing a page in llms.txt does not let it through.
What it does do is narrower and real. Writing one forces you to state, in one place and in plain language, what your site is and which twenty pages matter. Most teams have never actually done that. The result is a file you can hand to any AI tool, paste into a prompt, or point a partner at, and it costs an hour.
Who actually publishes one
The format has serious adopters even though no engine documents reading it. OpenAI, Anthropic and Google all publish an llms.txt for their own developer documentation, and Chrome’s Lighthouse has audited for one since version 13.3 moved its agentic-browsing category into the default report.
Read a real one rather than a toy example: our own file runs to 86 links across ten sections, and its opening summary is written as the sentence we would want quoted back.
llms.txt, robots.txt and sitemap.xml do different jobs
- robots.txt is permission. It tells each crawler what it may not fetch, and it is the only one of the three that actually controls AI crawler access. The robots.txt generator covers every documented AI crawler token and what blocking each one costs.
- sitemap.xml is inventory. Every URL you want indexed, with no indication of which ones matter or what they are about.
- llms.txt is editorial. A short list of what is worth reading, with a human sentence explaining each one.
What about llms-full.txt?
A companion convention, started by Mintlify and Anthropic for documentation sites and not part of the llmstxt.org specification, that holds the full text of your key pages in one file instead of links to them. It saves a model the fetches, but it duplicates content you already publish and goes stale the moment you edit a page. A confidently wrong copy of your pricing is worse than no file. Start with llms.txt.
What is a brand knowledge file?
A related idea with a different shape. Where llms.txt points at pages, a brand knowledge file states the facts themselves in a structured form: what you sell, what it costs, who you serve, the questions buyers ask and the answers you stand behind. It is the difference between handing someone a reading list and handing them the briefing.
It has no published specification, so treat it as a convention rather than a standard. It is useful for the same reason llms.txt is: one place where your own facts are written down unambiguously, which you control. The brand knowledge file generator builds one from your brand facts, pricing and FAQs.
Should you make one?
Yes, if you can spend an hour and will keep it current. The cost is a few bytes and the case for it is that it is cheap and possibly useful, not that it is a lever. Update it when the structure of your site changes rather than on a schedule: a stale file that links to pages which have moved is worse than none, because every link in it is a claim about your own site that you are inviting a model to trust.
Build one with the llms.txt generator, publish it at your domain root, then confirm the live URL returns plain text with the llms.txt checker. Publishing is the step that most often goes wrong, especially on managed WordPress hosting.
Frequently asked questions
What is llms.txt used for?
It gives a large language model a short description of your site and a curated, linked list of the pages you want it to read, so the model is not left to infer your site from whatever it happens to crawl. It is a guide to your content, not a control over it.
Does llms.txt actually work?
No major AI platform documents reading it, and Google’s Search Central documentation states that Google Search ignores it. Treat it as unproven. What it reliably gives you is one accurate, machine-readable description of your own site that you control.
Is llms.txt a ranking factor?
No. Nothing published by Google, OpenAI, Anthropic or Perplexity treats llms.txt as an input to ranking or citation. Structured data, crawlability and clear answer-first content are the parts of AI visibility with actual support behind them.
Is llms.txt the same as robots.txt?
No. robots.txt tells crawlers what they may not fetch; llms.txt suggests what is worth reading and what it means. One is a fence, the other is a guide. They do not replace each other, and llms.txt grants no permission a crawler does not already have.
Where does the llms.txt file go?
At the root of your domain, so it loads at yourdomain.com/llms.txt, which is the same place robots.txt lives. It must be served as plain text. A file in a subfolder, or one that returns an HTML page, does not count.
Does llms.txt stop AI companies training on my content?
No. It is a guide, not a permission system, and it carries no legal weight. Controlling AI crawler access is a robots.txt job, with per-bot rules for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest.
What goes in a brand knowledge file?
Your own facts in a structured form: what you sell, pricing, who you serve, the questions buyers ask and the answers you stand behind. Unlike llms.txt it has no published specification, so treat it as a useful convention rather than a standard.
Do I need both llms.txt and a brand knowledge file?
Not necessarily. llms.txt is the reading list and the brand knowledge file is the briefing. If you only do one, do llms.txt, because it has a published format and takes an hour. Add the knowledge file when your pricing or product facts are the thing people keep getting wrong about you.