The llms.txt file, in plain words
An llms.txt file is a short plain-text document at the root of a site that lists its most useful pages for a language model to read. It is a proposal rather than a standard, and it is cheap to publish.
An llms.txt file is a short plain-text document at the root of a site that lists its most useful pages for a language model to read. It is a proposal rather than a standard, and it is cheap to publish.
Short answer
An llms.txt file is a plain-text file placed at the root of a website that points language models at the pages worth reading, with a line of context for each. It is a community proposal, not an official standard, and no engine is obliged to read it. It costs little to publish and it cannot hurt a site.
The file is written in ordinary markdown and lives at the root of a domain, alongside the robots file. It opens with the name of the site, a sentence saying what the site is, and then a set of links grouped under simple headings.
Each link carries a short note explaining what that page covers. The notes are the point. A bare list of addresses tells a reader nothing, while a list where every entry says why it exists can be understood without fetching anything.
There is no schema to satisfy and no validator to pass. If a person can read the file and understand the shape of your site from it, the file is doing its job. That is the whole design goal behind the proposal.
Honestly, we do not know in every case. Some tools and assistants fetch it. Others ignore it completely. No major engine has committed publicly to reading it as a ranking or citation input, and none should be assumed to.
That uncertainty is worth stating plainly, because a good deal of writing about this file implies otherwise. It is a proposal that gained traction quickly. Traction is not adoption, and adoption is not obligation.
The honest case for publishing one is the cost. It is a single short file, written once, updated when your site structure changes. Set against a small chance of being read more accurately, that trade is easy to make.
An XML sitemap lists every page you want crawled, in a machine format, with no commentary. It exists so nothing is missed. Our page on sitemap practice covers how that file should be built and kept current.
An llms.txt file is the opposite instinct. It lists a small number of pages and explains them. It exists so the important material is found quickly, which is a different job from making sure everything is found at all.
The two do not compete and neither replaces the other. Keep the sitemap complete and automatic. Keep the llms file short and curated. A site with both has covered the complete reading and the guided reading.
Begin with the pages a stranger would need in order to describe your business correctly. For most small sites that is the homepage, the services or products, the service area, the pricing page if you have one, and the contact route.
Then add the material that answers real questions. Guides, policies and anything that states a fact about how you work. Leave out tag archives, thin category pages and anything you would not want quoted back to you.
It does not grant permission and it does not withhold it. Access to your content is governed by your robots file, your terms, and the law, none of which this file changes. Treat it strictly as a map rather than as a gate.
It also does not rescue a site that is hard to read. If your headings skip levels and your pages open with atmosphere, a map to those pages simply delivers a reader to the same difficulty. Fix the pages first.
And it will not place you in an answer. Nothing does that on request. The file makes your material easier to understand correctly, which is a real benefit and a modest one, described accurately rather than oversold.
A stale map is worse than none, because it points confidently at material that has moved. Tie the file to the same moment you update navigation. If a page earns a place in the menu, it probably belongs in the file.
Check the links the way you would check any others. A dead entry in this file will not be reported to you by anything, so it can sit broken for a year. Our page on broken links covers the routine.
Keep the file small enough that maintaining it stays trivial. The moment it becomes a chore it will stop being updated, and an abandoned file is the one outcome here that actually costs you something.
Publishing this file is the last item of a list, not the first. Ahead of it sit headings in order, definitions near the top, markup that names the page, honest claims and pages that load quickly on a phone.
Those earlier items are what determine whether your material can be understood and repeated at all. The file only helps a reader choose which of your pages to read, which matters once the pages themselves are worth reading.
If you want the order in full, the page on generative engine optimization lays out the whole sequence, with the structural work that carries most of the weight placed where it belongs.
No. It is a community proposal that has been adopted by some sites and tools. There is no standards body behind it and no engine is obliged to fetch or honor it. Publish it because it is cheap and clear, not because it is required.
At the root of your domain, so it sits at the same level as the robots file and resolves directly under your domain name. It should be served as plain text and should not require a login or a redirect to reach.
No. It is a map, not a permission setting. Crawling and usage are governed by your robots file, your terms of service and applicable law. Adding or removing entries here changes what is suggested, not what is allowed.
It may help a system understand your site correctly, which is a modest and real benefit. It will not place you in an answer. The structural work on your actual pages does far more, and should be done first.
15-day free trial. Card required. Cancel before day 15 and you pay nothing.