What Is llms.txt? Its Impact on AI Visibility, With Evidence

Does an llms.txt File Actually Boost AI Visibility? 300,000-Site Data and Google's Answer

This article evaluates what the llms.txt file promises, Google's official position, the results of an independent 300,000-domain study, and which site types it actually makes sense for — with sources.

Category: SEO#GEO#AI Search#Technical SEO
Summarize with ChatGPT

llms.txt is a proposed file, published at the site root (/llms.txt), that lists a site’s most important content in a markdown format meant to be easy for AI models to read. The short answer: today this file does not measurably improve AI visibility. Google states outright in its official documentation that it doesn’t use the file. An independent study covering 300,000 domains also found no correlation between having the file and getting cited by AI systems.

Despite that, the topic won’t go away. Many agencies selling GEO services have made llms.txt a standard deliverable. In this piece, we separate what the file promises, what the evidence actually shows, and when publishing one makes sense.

What Is llms.txt, and Where Did It Come From?

The proposal was published by Jeremy Howard, founder of Answer.AI, on September 3, 2024, at llmstxt.org. It’s a plain markdown document made up of the following sections:

  • A mandatory single H1 heading containing the site or project name.
  • A blockquote with a short summary of the site.
  • H2-headed sections listing links to important pages along with brief descriptions.
  • An “Optional” section for secondary links that can be skipped when the context window is tight.

In practice, many sites also publish a companion llms-full.txt variant that merges the entire body of documentation into a single markdown file.

What Problem Does It Promise to Solve?

The proposal’s rationale rests on a technically reasonable observation. Language models’ context windows are too narrow to process an entire website. Converting the clutter of menus, ads, and scripts in HTML into meaningful text is also an error-prone process. llms.txt proposes skipping that conversion entirely and handing the model a tidy “site map and summary” instead.

The problem framing lines up with what we described in how AI chatbots search: bots read pages from server-rendered HTML and break them into chunks. According to the proposal’s advocates, a clean markdown source could shorten that chain. But a solution being logical doesn’t mean engines actually use it. That’s exactly where the evidence question begins.

Does Google Use llms.txt?

No. Google’s document on succeeding in AI features, updated in December 2025, answers this directly: there’s no additional requirement for appearing in AI Overviews or AI Mode — you don’t need to create new machine-readable files, AI text files, or special markup. Instead, the document lists the same fundamentals it always has:

  1. Pages should be indexable and snippet-friendly.
  2. robots.txt and infrastructure should allow crawling.
  3. Content should be discoverable through internal links.
  4. Important content should be kept in text format.
  5. Merchant Center and Business Profile information should be verified.

Because Google owns both the search market and the AI Overviews surface, this statement alone is decisive: most of a typical Turkish marketing site’s AI-visibility traffic comes from systems that never read llms.txt at all.

What Does the Independent Data Show?

SE Ranking measured the file’s adoption rate and its effect on citations in a 300,000-domain study published on November 7, 2025. The study combines Spearman correlation, XGBoost regression, and SHAP analysis. Two results stand out:

  • Only 10.13% of the domains examined had an llms.txt file at all.
  • No meaningful relationship was found between having the file and citation frequency in large language models. The researchers reported that prediction accuracy actually improved once the llms.txt variable was removed from the model — meaning the file was producing noise, not signal.

The study’s own conclusion is balanced: the file doesn’t drive citations today, but because the technical risk and cost are low, it can be kept as a bet on the future. We pass along the first half of that conclusion as-is; the second half we evaluate by site type below.

Who Actually Uses It, and Where Does It Make Sense?

The file’s real use case today isn’t marketing sites — it’s developer documentation. Anthropic publishes its own llms.txt file for the Claude developer docs, and documentation sites for companies like Stripe and Cloudflare offer similar files. In that context the file serves a genuinely useful purpose: coding assistants and agentic tools loading a library’s documentation as context may prefer a tidy markdown source.

That distinction matters for the decision. If you’re a SaaS company shipping API documentation, llms.txt is a low-cost addition with a concrete use case. If your goal is being cited in ChatGPT or Google AI Mode answers, the same file doesn’t deliver a measurable contribution today.

The Other Side of the Coin: Blocking and Charging Bots

The llms.txt discussion rests on the assumption “let’s serve more content to AI.” On the publisher side, though, a movement in the opposite direction is growing — and that’s where the real strategic decision is being made.

According to Cloudflare’s August 2025 data, the crawl-to-referral ratio diverges dramatically by platform. In July 2025, Google crawled roughly 5.4 pages for every 1 visit it sent; for OpenAI that ratio was around 1,100 pages, and for Anthropic around 38,000 pages. Anthropic’s ratio has been dropping fast from the 287,000 level at the start of the year, but the overall picture is clear: AI bots crawl heavily and refer sparingly.

Based on this picture, Cloudflare announced a policy change on July 1, 2026. Starting September 15, 2026, on new domains, bots used for model training and agentic purposes will be blocked by default on pages that carry ads, while search bots will remain allowed. Hybrid crawlers that run both search and training functions through the same bot will be blocked once training is blocked. The company also operates a marketplace that lets sites charge for content access.

The practical takeaway for site owners: the decision about AI bots isn’t a single “block/allow” toggle. As detailed in our GEO article, search bots, training bots, and user agents serve separate purposes, and robots.txt and CDN rules should be managed with that distinction in mind. A site that wants visibility accidentally blocking search bots is a far more costly mistake than simply lacking an llms.txt file.

A Decision Framework by Site Type

Site typellms.txt decisionRationale
Developer documentation, API, libraryWorth tryingCoding assistants may use it when loading documentation context; cost and risk are low
Marketing site, corporate site, blogNot a priorityGoogle doesn’t use it; independent data shows no citation effect; effort is better spent on SSR HTML and content structure
News site, publisherBot access policy comes firstThe real decision is bot access and monetization; llms.txt doesn’t resolve that

Whichever decision you make, two conditions don’t change. First, if you do publish the file, it must be kept consistent and up to date with the visible site content. A poorly maintained llms.txt carries the same inconsistency risk as putting an answer in schema that doesn’t appear on the page. Second, the file never substitutes for crawlable HTML under any scenario.

Our Decision

We don’t publish an llms.txt file on this site. Our reasoning boils down to three points:

  1. None of the engines our traffic comes from have stated that they use the file. Google has explicitly stated that it doesn’t.
  2. Independent data shows no citation effect.
  3. We’re putting the same maintenance effort into work with a measurable payoff instead: server-rendered clean HTML, content structured to answer at the paragraph level, and schema consistent with visible content.

This decision isn’t permanent — it’s conditional. If a major engine officially announces llms.txt support, or if a transparent study shows a citation effect, we’ll update this article with a dated revision.

Bottom Line

llms.txt is a tidy format proposed for a real technical problem. But it doesn’t offer a proven win for AI visibility today. Google doesn’t use the file, the 300,000-domain dataset shows no citation effect, and adoption sits around 10%. Teams shipping developer documentation can try it at low cost; the priority for marketing sites hasn’t changed: crawlable content, clear answer structure, and correct bot-access decisions.

For how to measure AI-originated visits, see our review of the GA4 AI Assistant traffic channel; for the full visibility strategy, see the GEO and AEO guide.

Sources

Read this topic inside a learning path

You can read this post on its own, or continue through the guide section to follow related topics in a clearer order.

Open learning pathsSend feedback

Değerlendirme

Bu yazı ne kadar faydalıydı?

1 ile 5 arasında puanla

Değerlendirme

Bu yazı işine yaradı mı?

Tek tıkla puanlayabilirsin.

1 ile 5 arasında puanla