What Is llms.txt, and Do You Need One?
llms.txt is a proposed plain-text file at your site root listing the pages AI systems should read, in Markdown. The honest evidence: no major engine documents consuming it, and 97 percent of the files in a 137,000-domain log study received zero requests. Ship one only because it is cheap, never as a visibility lever.
On this page 6 sections
What is llms.txt?
llms.txt is a proposed standard from Jeremy Howard of Answer.AI, September 2024: a Markdown file at your site's root that gives AI systems a curated, machine-readable summary of the site, what it is, which pages matter, where the clean versions live. llms.txt is designed to help large language models spend a finite context window well, and many sites also publish a companion llms-full.txt carrying the whole corpus in one file, a convention that grew up around the proposal rather than part of it.
The idea is reasonable on its face. Assistants work from fetched web content, fetching is expensive, and a site-provided index could save everyone the crawl, which is why developer docs platforms and API references were the earliest adopters: a docs page maps naturally onto a curated link list. The question that matters is not whether the llms.txt standard is elegant; it is whether anything reads the file.
Does anything actually read it?
Mostly, no, and this is measured rather than argued. Ahrefs analysed server logs across 137,000 domains and found 97 percent of llms.txt files received zero requests. Of the requests that did arrive, around 1 percent came from AI retrieval bots. The largest single group, 21.7 percent, was SEO audit tools, and 77 percent of the bots fetching the file were not AI tools at all.
The absence goes to the top of the stack: OpenAI, Anthropic, Google and Perplexity each document their crawlers, and none documents consuming llms.txt, so none of ChatGPT, Claude or Gemini can be shown to read it at retrieval time. Google's own guidance says Google Search ignores the files, so one neither helps nor harms there. SE Ranking's 300,000-domain analysis found no effect on citation frequency. Adoption spread anyway, to 28 percent of domains in Ahrefs' study, which says more about how advice spreads than about the file.
Some of that growth is now automatic rather than chosen. Our own survey of 359 live WordPress sites in September 2026 found 18.6 percent of the 354 that could be checked serving an llms.txt, and 61 of those 66 files sat on a site running a detectable SEO plugin. The measurement sits in what WordPress sites actually serve to AI crawlers.
How does llms.txt compare to robots.txt and sitemap.xml?
The three files sound alike and do opposite jobs. A robots.txt file excludes: it tells compliant crawlers what not to fetch, and search engine optimization has leaned on it for three decades. A sitemap.xml enumerates: every URL, no judgement, so nothing gets missed. An llms.txt file curates: the pages worth an AI system's attention, with a short description each. Exclusion, enumeration, curation.
The differences that matter are enforcement and adoption. Standards like robots.txt work because consumers exist: crawlers honour robots.txt imperfectly but measurably, and sitemaps demonstrably help search engines find URLs. The existing standards earned their place; llms.txt has a syntax and a website but, so far, no documented consumer among the major engines or AI models. Structured data sits in the same drawer: like schema markup, an llms.txt file can only mirror content that must already stand on its own, and neither has been shown to buy AI citations.
From RankX AIWebsite AuditFind what keeps your pages out of AI answers.Explore Website AuditWhy publish one anyway?
Because the cost can be zero, and at zero cost even a small option is worth holding. This site serves an llms.txt and a full llms-full.txt corpus, generated from the same constants and collections that render the visible pages, so the file cannot drift from the site and its maintenance cost after setup is nothing. If an engine starts consuming the convention, we are already legible to it; if none ever does, we spent nothing that mattered. Sites that use llms.txt today do it for that option value, not for a measured effect.
The same logic in reverse is a warning sign worth naming: an SEO expert whose audit leads with your missing llms.txt is leading with the cheapest, least consequential box on the checklist. The retrieval mechanics that decide visibility, rendering, crawler access, extractable structure, are covered in the GEO guide; none of them lives in this file.
How do you create an llms.txt file?
The basic structure comes straight from the llms.txt proposal, and any Markdown tool can parse it: one h1 heading with the site name, a blockquote carrying a concise summary, then h2 headers for each section, Home, Docs, Blog, whatever fits, each holding a list of URLs with a short description per link. Save it as a plain text file, upload it to the site root so it resolves at /llms.txt, and serve it as text; no plugin, no build step, one file in a public directory.
The free llms.txt Generator reads your site and drafts the file: it finds the pages worth listing, writes the descriptions from what the pages actually say, and outputs the structured format ready to serve at the root. It runs without an account. Whichever way you create llms.txt content, four rules keep it useful:
- Lead with what the site is, in one paragraph a machine can quote.
- Curate the important, high-value content that answers real questions, not every URL; the sitemap already handles exhaustive.
- Write honest one-line descriptions of your best content. The file's only conceivable reader is a system deciding what to fetch; a description that oversells earns a fetch that disappoints.
- Generate it from your content source if you can, so it stays up-to-date content rather than a snapshot. A hand-written file is out of date at the first publish after it.
The adjacent idea with more substance: Markdown twins
A related convention has real infrastructure behind it: serving Markdown versions of each page, either at a .md URL or through an Accept header, which Cloudflare now supports at the edge. This site does that too, every HTML page has a Markdown twin generated from the same read, with the HTML kept as the indexable surface. The honest caveat is the same shape as before: no engine has stated it requests Markdown. Engines already convert your HTML to Markdown themselves, which is the actual lesson of the whole area: the durable investment is clean, semantic, server-rendered HTML with nothing meaningful hidden behind JavaScript, because that is the input every pipeline shares. Whether your pages pass that bar is checkable with the site audit in Website Audit.
Questions about GEO Fundamentals
Will llms.txt improve my rankings or AI citations?
No measured effect exists. SE Ranking's 300,000-domain study found no effect on how often a domain is cited, Google states that Google Search ignores the file, and no major engine documents consuming it. Anyone selling llms.txt as a visibility lever is selling ahead of the evidence.
Is llms.txt mandatory?
No. No search engine, AI vendor or web standards body requires it, Google states that Google Search ignores it, and nothing breaks without one. Publish an llms.txt file only as a cheap option on a future where some AI system starts consuming the convention, nothing more. RankX AI's free llms.txt generator writes one from your sitemap in seconds, so the option costs almost nothing.
Should llms.txt list every page on the site?
No. The proposal is a curated index, not a sitemap dump: the pages that answer real questions, each with a one-line honest description. A file that lists everything ranks nothing first, and the sitemap already exists for exhaustiveness. Curation is the only editorial decision the format actually asks of you.
Related reading
Generative Engine Optimization: A Complete Guide
What Generative Engine Optimization is, how generative engines retrieve and cite pages, how GEO extends SEO, what measurably works and what failed testing.
How AI Crawlers Read Your Site
Which AI crawlers exist, their user agents, what they fetch, which of them run JavaScript and which do not, and how to verify your pages are reachable.
Sources
- Ahrefs, llms.txt server-log study, 137K domains: 97 percent of files never requested, 77 percent of fetching bots not AI tools, SEO audit tools 21.7 percent (opens in a new tab) Checked 1 Oct 2026.
- llms.txt proposal, Answer.AI (opens in a new tab) Checked 1 Oct 2026.
- SE Ranking, LLMs.txt and AI citations across nearly 300,000 domains, 7 November 2025: no effect on citation frequency (opens in a new tab) Checked 1 Oct 2026.
- Google Search Central, optimizing for generative AI features: Google Search ignores LLMS.txt files (opens in a new tab) Checked 1 Oct 2026.
- OpenAI crawler documentation (no llms.txt consumer documented) (opens in a new tab) Checked 1 Oct 2026.
- Anthropic crawler documentation (no llms.txt consumer documented) (opens in a new tab) Checked 1 Oct 2026.
- Perplexity crawler documentation (no llms.txt consumer documented) (opens in a new tab) Checked 1 Oct 2026.
CoversThis article covers the GEO Fundamentals topic, the Website Audit feature and the llms.txt Generator tool.
Terms usedllms.txt, AI Crawler, Robots.txt, GPTBot and XML Sitemap.
Read this page asMarkdown: /blog/llms-txt.md.
All articlesEverything RankX AI publishes is listed on the blog index.
Ask an assistantAsk ChatGPT (opens in a new tab), Ask Claude (opens in a new tab) or Ask Perplexity (opens in a new tab).
Preferred sourceIf Google is your front door, you can add RankX AI as a preferred source (opens in a new tab), which asks your own results to surface more of what we publish.
