An llms.txt file is a plain Markdown file, usually placed at a site’s root, that gives AI systems a short summary of a site plus links to its key pages. The specification is real and well-defined, but it is a community proposal, not a web standard, and Google has stated on the record that its crawlers do not request or use it. Whether publishing one is worth your time depends on what you actually want it to do.
What the llms.txt specification actually requires
The llms.txt specification, proposed by Jeremy Howard in 2024, defines a precise Markdown structure rather than a vague “AI-friendly text file.” Most sites that publish one get the structure wrong, which defeats the point: the file exists so that automated systems can parse it with “fixed processing methods,” and a file that breaks the order breaks that parsing.
| Element | Required? | Rule from the spec |
|---|---|---|
| H1 title | Required — the only mandatory section | Must be the first line, a single Markdown H1 |
| Blockquote summary | Optional but recommended | A short blockquote with the project’s key context, placed right after the H1 |
| Free-text details | Optional | Any Markdown except further headings, used for background an agent needs before the links |
| H2 file-list sections | Optional, zero or more | Each is a Markdown list of [name](url): notes links grouped under an H2 |
| “Optional” section | Convention, not a rule | By convention, an H2 literally named “Optional” holds links an agent can skip under a tight context budget |
Two details trip up most implementations. First, only the H1 is mandatory — a one-line file with just a title is technically valid, even if it’s not useful. Second, the free-text section cannot contain further headings; once a heading appears, the spec treats everything after it as a file-list section, so burying prose under an H2 silently turns it into a broken link list.
A minimal file, rewritten to match the spec
Here is the kind of file most teams ship first, and why it doesn’t actually parse the way they expect.
Before — looks reasonable, fails the spec:
# Acme Docs
Welcome to Acme. We make project management software.
## Getting Started
Read our quickstart guide at acme.com/docs/start
## Pricing
See plans at acme.com/pricing
- The second line isn’t a blockquote, so an agent parsing strictly for the spec’s blockquote summary finds nothing.
- “Read our quickstart guide at acme.com/docs/start” is prose with a bare URL, not a Markdown link — the spec requires
[name](url)inside list items, so this line doesn’t register as a file entry at all.
After — matches the spec’s parsing rules:
# Acme Docs
> Acme is project management software for distributed teams; this file indexes our docs and pricing for automated tools.
## Docs
- [Quickstart guide](https://acme.com/docs/start): environment setup and first project
## Optional
- [Pricing](https://acme.com/pricing): current plan tiers
The fix is mechanical: turn the summary into an actual blockquote, and turn every reference into a real Markdown hyperlink inside a list item. An agent built against the spec can now extract the summary and the two links reliably; it couldn’t extract either from the first version.
What Google has actually said about it
Google’s public position is unusually direct for an SEO topic. Google Search Advocate John Mueller has compared llms.txt to the old keywords meta tag, stating that no AI services used it and that bots simply didn’t request the file, according to reporting from Search Engine Journal. At a Google Search Central Live session in the Asia-Pacific region, Gary Illyes and Amir Taboul reportedly confirmed Google is not pursuing llms.txt as a crawling input.
That statement is specifically about whether Googlebot or Google’s AI features fetch and act on the file — it is not a claim that structured Markdown summaries are worthless in general. Google’s own crawling and indexing documentation still describes the signals it does use: robots.txt for crawl permissions, sitemaps for discovery, and standard HTML for content extraction. llms.txt sits outside all three.
Which crawlers, if any, actually request the file
Public, verifiable evidence that a major AI provider’s crawler fetches llms.txt by default is thin. The file’s own spec site documents the convention and lists tools that generate or consume it, but adoption by the crawlers behind ChatGPT, Claude, Gemini, or Perplexity has not been confirmed with server-log evidence from those providers themselves. Treat any site claiming a traffic or citation lift “because of llms.txt” as an unverified correlation unless it shows request logs, since the same period usually includes other content and link changes too.
How llms.txt relates to robots.txt and your sitemap
These three files answer different questions, and conflating them is the second most common mistake after getting the Markdown structure wrong.
| File | Question it answers | Who actually reads it |
|---|---|---|
| robots.txt | Which URLs may a crawler fetch at all? | Googlebot, Bingbot, and other declared crawlers that honor the Robots Exclusion Protocol |
| sitemap.xml | Which URLs exist, and when did they last change? | Search engines, for discovery and recrawl scheduling |
| llms.txt | What is this site about, and which pages matter most? | Tools and plugins that implement the convention; confirmed not read by Google’s crawlers |
None of the three substitutes for another. A site can have a perfect robots.txt and sitemap.xml and still ship a broken llms.txt, and fixing the llms.txt file won’t change how robots.txt or the sitemap are processed.
Generating one without hand-writing Markdown
For a WordPress site specifically, hand-maintaining this file is rarely necessary. Both Yoast SEO and AIOSEO ship llms.txt generation as a plugin feature, producing output that follows the spec’s structure automatically from your existing post and page data, per the tooling list on the llms.txt specification site. Documentation platforms like Mintlify and GitBook do the same for docs sites, and VitePress and Docusaurus have dedicated plugins for the same job. If you’re already running an SEO plugin that offers this, turning it on and then checking the output against the structural rules above is faster and less error-prone than writing the file by hand.
A checklist before you publish one
If you decide to ship an llms.txt file anyway — for internal tooling, a RAG pipeline, or on the chance adoption grows — these are the mistakes worth checking for first.
- Put it at the true root.
/llms.txt, not inside a subfolder, unless you intend it to scope only that subfolder — the spec explicitly supports path-scoped files, and the most specific one wins when several exist. - Make the first line a real H1. A plain line of text with no
#is not a title under the spec; the parser has nothing to anchor on. - Use an actual blockquote for the summary. A paragraph that merely sounds like a summary doesn’t satisfy the structural rule — it needs the
>marker. - Keep headings out of the free-text section. The moment a heading appears, everything after it is parsed as a file list, so a stray H3 in your background text will silently swallow your prose into the wrong section type.
- Use real Markdown links in file lists, not bare URLs.
[Quickstart](url)parses; “see our quickstart at url” does not. - Don’t duplicate your sitemap. llms.txt is meant to be a short, curated index — not an auto-generated mirror of every URL in sitemap.xml — otherwise it provides no summarization value over the sitemap itself.
- Keep it current. A stale file linking to pages you’ve since removed or merged is worse than no file, since any agent that does parse it will cite the dead link.
Most of this overlaps with ordinary technical SEO hygiene: it’s the same discipline behind a clean robots.txt versus meta robots setup, and the same instinct that drives a scaled content abuse audit — keep what you publish accurate, current, and structured the way the consuming system expects.
Does this replace anything you already have?
No. llms.txt does not replace robots.txt, which still governs crawl permissions, and it does not replace your XML sitemap, which still governs URL discovery for search engines. It also doesn’t substitute for the content quality work that actually earns citations in AI answers — the kind of page-level clarity covered in a thin-content audit or in matching pages to real search intent does more for visibility than an index file that the largest search engine says it ignores.
FAQ
Should I still publish an llms.txt file if Google says it doesn’t use it?
It costs little to maintain a short, accurate one once your content is already organized, and other AI products or internal tools may consume it even where Google’s crawlers don’t. Just don’t expect it to move search or AI-answer visibility on its own, and don’t let maintaining it take time away from higher-leverage work like structuring FAQ content for featured snippets.
Will an incorrectly formatted llms.txt file hurt my SEO?
There’s no evidence it affects search rankings either way, since Google has stated its systems don’t parse it. The only real cost of getting the format wrong is that any tool that does try to read it will fail silently, which defeats the reason to publish one in the first place.



