On June 17, 2025, Google search advocate John Mueller answered a question about llms.txt on Bluesky with a single line that site owners have been quoting ever since. "FWIW no AI system currently uses llms.txt," he wrote. In a follow-up reported by Search Engine Roundtable he added that "the consumer LLMs / chatbots (the ones that SEOs want traffic from) will fetch your pages - for training and grounding, but none of them fetch the llms.txt file."

Eleven months later, Chrome's Lighthouse audit tool began checking websites for an llms.txt file inside a new Agentic Browsing category, Search Engine Land reported on May 20. Asked about the apparent contradiction, Mueller said, "The short answer is that it's not done for search. There's more to websites than just SEO."

For a solo webmaster deciding what to put in the root folder this fall, the practical answer has settled. An llms.txt file is cheap to publish and does nothing for Google rankings. The controls that decide whether AI companies can train on your pages, cite them or fetch them on demand live in robots.txt and in your CDN dashboard, and one of those dashboards changed its defaults on September 15.

Where llms.txt Stands in September 2026

Jeremy Howard of Answer.AI published the proposal on September 3, 2024, describing it as "a proposal to standardize on using an /llms.txt file to provide information to help agents use a website." The file is plain Markdown at the site root, meant to hand AI tools a clean summary and a map of the pages worth reading. On August 10 of this year, Howard posted a revised version, noting on llmstxt.org that "this is v2 of the proposal, updated based on what I learned from two years of adoption."

Adoption has grown from a rounding error to a visible minority. A September scan by Rankability, run against the Tranco list of popular domains dated September 17, found an llms.txt file on 9.3 percent of the top 1,000 domains and 8.3 percent of the top 10,000. Only 1.6 percent of the top 1,000 also published the longer llms-full.txt variant. Earlier readings used different methods, so the numbers work as a snapshot rather than a clean growth curve.

Google's documentation, updated July 10, now addresses the question directly. "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them," reads its guide to AI features. The same page adds that "it's completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files."

The labs themselves publish llms.txt files for their own developer documentation, the proposal site notes, naming OpenAI, Anthropic and Gemini. Publishing a file is a different act from reading one, and none of the three has said in its crawler documentation that its bots consult llms.txt when deciding what to fetch. Our May guide to getting your site cited by AI chatbots treated the file as an early-mover bet, and the bet still costs about ten minutes of writing.

The Crawlers That Actually Arrive, and What Each One Is For

The real controls sit in robots.txt, and the major vendors now split their bots by purpose, which lets a site owner say yes to one use and no to another. OpenAI's bot documentation lists GPTBot to "crawl content that may be used in training our generative AI foundation models" and OAI-SearchBot to "surface websites in search results in ChatGPT's search features." A third agent, ChatGPT-User, handles fetches a person triggers inside a conversation, and the company notes that "because these actions are initiated by a user, robots.txt rules may not apply."

Anthropic's help center article, last updated April 7, describes three agents. Blocking ClaudeBot "signals that the site's future materials should be excluded from our AI model training datasets." Claude-User retrieves pages when a person asks Claude a question, and blocking Claude-SearchBot "may reduce your site's visibility and accuracy in user search results." The company also honors the Crawl-delay directive.

Google handles training through a token rather than a separate crawler. Google-Extended controls whether content is used "for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding," according to its crawler documentation. The token "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." AI Overviews run on ordinary Googlebot crawling, so blocking Google-Extended does not remove a page from them.

The Rest of the List

Perplexity states in its crawler documentation that PerplexityBot is "designed to surface and link websites in search results on Perplexity" and "is not used to crawl content for AI foundation models." Its user-triggered agent, Perplexity-User, "generally ignores robots.txt rules." Apple's Applebot-Extended, Apple explains, "does not crawl webpages" at all and only governs whether content already gathered by Applebot trains Apple's models.

Common Crawl's CCBot builds an open archive that many model builders draw on, and the nonprofit publishes its IP ranges and warns that impostors use its name. A small site that objects to training generally blocks GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended, and leaves OAI-SearchBot, Claude-SearchBot and PerplexityBot open so its pages can still appear as cited sources in AI answers.

Robots.txt remains a request rather than a lock. TollBit's first-half 2026 State of the Bots report found that about 15 percent of AI page fetchers reached disallowed URLs on European sites. ChatGPT-User, Bytespider and Youbot each did so on nearly half of the European sites that had named them, Search Engine Journal reported on August 14.

Cloudflare Changed the Defaults on September 15

For the roughly one fifth of the web that runs behind Cloudflare, the network edge has become the enforcement layer. On July 1, 2025 the company announced that it would block AI crawlers by default on new domains. It also introduced pay per crawl, a private beta that uses the long-dormant HTTP 402 Payment Required status to let owners allow, charge or block each crawler.

Its measurements explain why publishers asked for those tools. In July 2025, Cloudflare counted about 38,000 pages crawled by Anthropic for every referral sent back, 1,091 for OpenAI, 195 for Perplexity, 41 for Microsoft and 5 for Google. Training accounted for 79 percent of AI crawling in that period.

This summer the company went further. Starting September 15, 2026, according to a July 1 post, new domain registrations have training and agent crawlers blocked by default on pages that display ads, while search crawlers remain allowed. Cloudflare now sorts bots into Search, Agent and Training groups, and its managed robots.txt appends a use=reference signal. TechCrunch's account describes a broader reach that includes free-tier users, so owners on existing accounts should open the AI Crawl Control panel rather than assume which rule applies to them.

The Standards Still Being Written

Two efforts aim to replace the patchwork of user agents with a single vocabulary. The IETF's AI Preferences working group published draft-ietf-aipref-vocab-08 on September 14. It remains an Internet-Draft with sections the document itself says do "not yet have consensus." Really Simple Licensing, released as RSL 1.0 on December 10, 2025, adds machine-readable licensing terms to robots.txt. Its backers list more than 1,500 supporting organizations, including Cloudflare, Akamai, Creative Commons and the Associated Press.

Neither standard changes what a small site should do this month, because the crawlers that matter today still read the user-agent rules they have documented. Our June look at AI-compliant hosting covered the firewall layer, and robots.txt remains the first file every compliant crawler requests.

A Configuration for a Small Site This Fall

Start with a decision about training, since it is the one choice with no traffic attached. A site that objects adds separate Disallow rules for GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended. A site that welcomes training leaves them alone, knowing that the Cloudflare figures show training crawlers sending little or no traffic back.

Keep the search agents open unless you have a specific reason to close them. OAI-SearchBot, Claude-SearchBot and PerplexityBot are how pages become cited sources in ChatGPT search, Claude and Perplexity answers, and the vendors say those bots do not feed model training. This is the core of an AI crawler robots.txt strategy that protects content without disappearing from answer engines.

Publish an llms.txt if your site has documentation, product pages or tutorials that an AI agent might need to navigate, and skip it without worry if your site is a local business brochure. Then check your CDN or host's bot controls, because a Cloudflare AI crawler default set at registration can override the intentions you wrote into robots.txt. Once a quarter, review your server logs for user agents that fetch pages you told them to leave alone.

Mueller's June 2025 post ended on a question about whether chatbots might start fetching llms.txt, followed by a joke about winning the lottery. The Lighthouse audit shipped eleven months later.