Skip to content
RankX AI

Free tool

The llms.txt Generator writes a spec-valid file, not a dump.

RankX AI reads your sitemap, picks the pages that matter, writes their notes from your own metadata, and validates the result against the llms.txt specification before you see it.

Free. No account. We read your sitemap and up to 60 pages, one request at a time, honouring your robots.txt. Results cache for 24 hours.

Free AI audit

What it is

Generate llms.txt in about a minute, free and checked against the spec.

llms.txt is a proposed markdown file at your site root that lists your most useful pages, with a short note each, for AI agents to read.

  • What it is

    A short reading list, written in markdown

    An llms.txt file names your site in one H1, sums it up in a blockquote, then lists your key pages under H2 headings, each a link with an optional note. Agents read it on demand to find the right page. The format is a community proposal, now in its second version, not a standard.

    It points AI systems at pages; it changes nothing about what a crawler may fetch.

  • Who it is for

    Documentation sites first, everyone else optionally

    The proposal's own author says llms.txt files are used most heavily for software documentation, where coding agents follow them to API references and tutorials. A marketing or SEO team does not need an llms.txt, but can publish one cheaply as a clean map of its best pages, provided nobody expects it to lift rankings or AI citations.

    If the goal is being cited by ChatGPT or Google, crawler access and the page itself decide that.

  • What it costs to run

    A free llms.txt generator, with no account

    The llms.txt Generator is free and needs no account: enter your website URL and it generates an llms.txt file in about a minute. A run considers up to 100 URLs from your sitemap, fetches at most 60 pages one request at a time while honouring your robots.txt, keeps about 50 links and caches the file for 24 hours. Notes come from your meta descriptions, so a page without one gets a bare link.

    No sitemap means a file built from the homepage alone, and a thinner one.

What this generates

A valid llms.txt: small, ordered, stricter than most tools respect.

The llms.txt Generator follows the published grammar: one required element, a short structure, validated output. Most generators emit a sitemap dump in markdown clothing.

The grammar, in five rules

  1. 1. One H1 naming the site. It is the only required element in the whole format.
  2. 2. An optional blockquote directly after it, carrying a short summary a model should read first.
  3. 3. Optional prose next: paragraphs or lists, with no headings of any level.
  4. 4. H2 sections last, each a markdown list whose entries are a link, name and URL, with an optional colon and note.
  5. 5. An H2 named Optional, by convention, holds secondary links an agent can skip. Since v2 it is a convention only, with no mechanical meaning.

Files can also live at subpaths such as /docs/llms.txt. Version 2 of the proposal, published in August 2026, says a file covers the pages under its path and that agents should use the most specific one, and it adds a way for a page to point at the file that covers it: a link with rel="describedby".

Mechanically it is a plain text file, like robots.txt, and it is not a sitemap.xml replacement: a sitemap lists every URL for traditional search engines to crawl, while llms.txt hands AI models a short curated reading order with none of your navigation, cookie banners or footer chrome around it. That is the whole appeal to anyone building on assistants like ChatGPT, Claude or Gemini, and the whole reason to create one by hand rather than dumping a URL list.

An llms.txt example, as the generator writes it

llms.txtExample
# Harbour Joinery

> Harbour Joinery designs, makes and fits bespoke kitchens and fitted furniture across Bristol and Bath.

Generated by the RankX AI llms.txt Generator on 2026-10-01 from harbourjoinery.test's sitemap. Curated to the most important pages; edit freely before publishing at /llms.txt.

## Key pages

- [About us](https://harbourjoinery.test/about): A workshop of six joiners in Bristol, making kitchens and furniture since 2009.
- [Contact](https://harbourjoinery.test/contact): Book a home visit or ask for a written quote.

## Services

- [Bespoke kitchens](https://harbourjoinery.test/services/bespoke-kitchens): Hand-made kitchens in oak, walnut and painted finishes, designed and fitted in eight to twelve weeks.
- [Fitted wardrobes](https://harbourjoinery.test/services/fitted-wardrobes): Built-in wardrobes and alcove units made to measure for older houses.

## Pricing

- [What a kitchen costs](https://harbourjoinery.test/pricing): Typical project ranges, the deposit and what every quote includes.

## Blog

- [How long does a bespoke kitchen take?](https://harbourjoinery.test/blog/kitchen-timeline): The design, making and fitting stages, week by week.

## Optional

- [Careers](https://harbourjoinery.test/careers)
- [Privacy policy](https://harbourjoinery.test/privacy)

A fictional business on a reserved .test domain. Grammar per llmstxt.org, v2, August 2026, verified 1 October 2026.

What the evidence shows

Almost nobody reads these files yet; you deserve to know.

Competing generators sell llms.txt as an AI SEO tactic. The measured evidence says it is not, and an AI search company owes you the measurements.

97% of files receive zero requests

Ahrefs studied 137,210 domains in 2026. About 28% publish an llms.txt, and 97% of those files got no requests at all in May 2026. Of the few that were fetched, the largest share of requests, about 22%, came from SEO audit tools, and AI retrieval bots made 1.1%. No AI bot asked for a file that did not exist. Publishing the file is cheap; expecting traffic from it is not supported by anything measured.

Google says it plainly

Google's guide to generative AI features, last updated 10 July 2026, lists llms.txt files among the things you can ignore: you do not need AI text files to appear in Google Search, including its AI features, because Google Search itself does not use them. OpenAI, Anthropic and Google all publish llms.txt files for their own developer documentation, and none documents reading yours.

The fabricated directive to avoid

A claim circulates that Anthropic honours a Commercial-Use directive inside llms.txt, sometimes with a whole invented field set around it. It appears in no version of the specification and no vendor documentation, and AI search summaries now repeat it as fact. If you want to control training use of your content, robots.txt tokens are the documented mechanism, and our crawler checker tests them.

Publishing it

Serve it at the root, then treat it like content.

Generate llms.txt above from your own sitemap and metadata; three habits keep the file worth publishing.

Put it at /llms.txt

Host it at your domain root as a plain text file, beside robots.txt. A documentation section that deserves its own index can carry a second file at /docs/llms.txt, which covers the pages under that path, and the most specific file wins.

Edit before you publish

The generator ranks pages by sitemap structure and depth, which is a good guess and only a guess. You know which ten pages actually explain your business; promote them, cut the rest, and keep the file under about 50 links so it stays small enough to read whole.

Update it when pages change

A stale index is worse than none, because the one reader it might get is being pointed at the wrong pages. Keep the file in version control with your site, regenerate after significant content changes, and re-run this tool any time; a repeat run inside a day returns the cached result instantly.

llms.txt beside robots.txt and sitemap.xml

robots.txtsitemap.xmlllms.txt
Its jobSays which paths a crawler may fetchLists every URL you want indexedPoints agents at your best pages, with notes
FormatPlain text rules per user agentXML, one entry per URLMarkdown: an H1, a summary, lists of links
StatusInternet standard, RFC 9309Sitemaps protocol, read by Google and BingCommunity proposal, v2 since August 2026
Who reads itSearch and AI crawlers, before they fetchSearch engine crawlersCoding agents and LLM tools, on demand; no AI search engine documents it
Where it lives/robots.txt, one per hostAny URL, declared in robots.txt/llms.txt, or a subpath it then covers

Keep going

One file moves nothing alone: these checks measure what does.

AI visibility is decided by crawler access, page mechanics and what Google's AI surfaces cite. Each has a free check here.

What decides whether an assistant can read you

Two things, and llms.txt is neither: whether an AI assistant's crawler is let in at all, which the AI Crawler Access Checker tests for every documented bot, and whether your words are in the HTML without JavaScript, which the AI Readiness Score measures page by page.

In RankX AI, the daily site signal watch re-reads your llms.txt, your robots.txt and your homepage every day and scores the llms.txt it finds, and on WordPress the RankX AI WordPress integration can serve an llms.txt from your site's root without a file on disk.

Questions

What people ask before publishing an llms.txt.

Direct answers, with the evidence behind each one.

What is an llms.txt generator?

An llms.txt generator is a tool that writes llms.txt, the proposed markdown file at a website’s root that lists its most useful pages for large language models and AI agents. The RankX AI llms.txt Generator reads your sitemap, keeps about fifty pages, writes each note from your own meta descriptions and checks the finished file against the spec.

Where did llms.txt come from, and what is in the file?

llms.txt, named for the large language models (LLMs) it serves, was proposed by Jeremy Howard of Answer.AI in September 2024 and revised as version 2 in August 2026. The file is deliberately small: one H1 with the site name, which is the only required element, an optional blockquote summary, optional prose, and then H2 sections whose entries are markdown links with optional notes. It is a suggestion to AI agents, not a control, and no AI search engines document reading it. The longer story is in what llms.txt is, and whether you need one.

Does llms.txt actually work?

Not for search visibility, on any evidence published so far. Ahrefs studied 137,210 domains in 2026: 28% publish an llms.txt, and 97% of those files received no requests at all in May 2026. Of the requests that did arrive, the largest share, about 22%, came from SEO audit tools, and AI retrieval bots made 1.1%. No AI bot asked for a file that did not exist. Where llms.txt does work is documentation, where coding agents follow it to the right page.

Does Google use llms.txt?

No. Google’s guide to generative AI features in Search, last updated 10 July 2026, lists llms.txt files among things you can ignore and says you do not need AI text files to appear in Google Search, including its generative AI features, because Google Search itself does not use them. Google’s John Mueller had earlier compared llms.txt to the keywords meta tag. Generate one for the cheap upside if you want it, but expect nothing from Google rankings or Google AI Overviews.

Why offer an llms.txt generator at all, then?

Because the demand is real, the cost is nothing, and an honest generator beats a hyped one. More than a quarter of studied domains publish the file, documentation platforms build one automatically, and the file doubles as a clean, human-readable map of your best pages that coding agents and documentation tools already use. What RankX AI will not do is sell llms.txt as an AI visibility tactic, because the evidence says it is not one. The generator exists so that if you want to generate the file, you get a correct, curated one in about a minute.

What makes an llms.txt file valid?

The spec is stricter than most generators respect. A valid file has exactly one H1 naming the site, and that H1 is the only required block. After it come, in order, an optional blockquote summary, optional prose with no further headings, and then H2 sections containing markdown lists where each entry is a link in the form of a name and URL, optionally followed by a colon and a note. No H3s anywhere, and each H2 section exists to hold its list of links. A built-in llms.txt validator checks this generator’s own output against those rules before handing it to you.

Does the Optional section still matter?

Less than it did. An H2 named Optional marks secondary links an agent can skip when it needs a shorter context, and the llms.txt Generator puts its overflow there. Version 2 of the proposal, published in August 2026, keeps Optional as a useful convention but removes its mechanical meaning, because the tool that once expanded a file into context and omitted that section is no longer part of the proposal.

How do I create a full llms.txt file?

A full file usually means llms-full.txt, a documentation-platform convention rather than part of the llms.txt specification, which never mentions it. The difference between llms.txt and llms-full.txt is size: platforms such as Mintlify and Fumadocs build the full one by concatenating every documentation page into one markdown file, so an agent can ingest everything in a single fetch. To create one yourself, export each page as markdown and join them in order. It suits docs sites, not marketing sites, so this tool generates llms.txt only and is not an llms-full.txt generator.

Is there a Commercial-Use directive for llms.txt?

No. A claim circulates that Anthropic honours a Commercial-Use directive inside llms.txt, sometimes expanded into a whole invented directive set with Training-Data and Citation-Required fields. It appears in no version of the specification and in no vendor documentation from Anthropic or anyone else, and it traces back to a single SEO blog post that cites no source. Be careful: AI search summaries repeat it as fact. If you want to control training use, robots.txt tokens are the documented mechanism.

Why does the generator cap the file at around 50 links?

Because curation is the entire point of the format. An llms.txt that mirrors your whole sitemap is just a worse sitemap, and the specification asks for a file small enough to fit in an agent’s context, with the detail behind the links. The generator ranks pages, keeps roughly the best fifty, orders sections commercial pages first and blog last, and writes each link’s note from the page’s own meta description. Dumping every URL is the measured weakness of most competing generators.

How does the generator build the file?

The generator reads your sitemap, or discovers it from robots.txt, considers up to 100 of its URLs and fetches at most 60 pages, one request at a time, while honouring your robots.txt. It extracts each page’s title and meta description, groups pages into sections, ranks them, and writes a spec-valid file with your homepage’s description as the blockquote summary. With no sitemap it falls back to your homepage alone. You can edit anything before you publish it, and a repeat run inside 24 hours returns the cached result instantly.

Where do I put the file once I have it?

Serve it at the root of your site as /llms.txt, next to robots.txt. The specification also allows a file at a subpath, such as /docs/llms.txt, which then covers the pages under that path, and agents should use the most specific file that applies. Serve it as plain text or markdown, keep it in version control like any other content, and update it when your key pages change.

How is llms.txt different from robots.txt and sitemap.xml?

They answer three different questions. robots.txt says what crawlers may fetch, is an internet standard, RFC 9309, and the major AI vendors document how their bots treat it. sitemap.xml lists every URL you want search engines to index. llms.txt points AI agents at your best pages first, is a proposal with no committed consumers, and enforces nothing. A site can sensibly have all three. The AI Crawler Access Checker tests the first; this generator writes the third.

Can you show an example of an llms.txt file?

Yes, and the page above carries a full one. The smallest valid llms.txt is a single line, an H1 such as "# Example Site". A useful one adds a blockquote summary on the next line, then H2 sections such as "## Services", each holding a list whose items are a page name in square brackets, its URL in round brackets, then a colon and a short note, and ends with an "## Optional" section for pages an agent can skip.

How often should I update my llms.txt file?

Update your llms.txt file whenever the pages it lists change: a new service, a retired product, a moved URL. A stale llms.txt points the few agents that read it at the wrong pages, which is worse than no file at all. Keep it in version control with the site, regenerate after significant content changes, and check it quarterly otherwise. In RankX AI, the daily site signal watch re-reads the file and records when it disappears or its quality score drops.

Start here

See where you show up in AI answers today.

Add your site and RankX AI suggests prompts, tracks your keywords and audits your pages, with first results minutes after setup.

7-day free trial. No credit card required. Cancel anytime.