Home Knowledge base llms.txt, one year on: what the server logs say
Research

llms.txt, one year on: what the server logs say

Two large studies in 2025 and 2026 checked whether AI crawlers actually request /llms.txt and whether having one correlates with being cited. The answers are 'almost never' and 'no'. What the file is still good for, and why the audit keeps checking it.

/llms.txt does not get you cited. Ahrefs checked the server logs of 137,210 domains in May 2026 and found that 97% of published llms.txt files received zero requests in the month. SE Ranking modeled AI citation frequency across roughly 300,000 domains and found no correlation between having the file and being cited; removing the variable made their model more accurate. Google has said it does not read the file and does not plan to.

We still check for it, and this article explains why, but the honest summary is that llms.txt is an agent-readiness signal, not a citation signal.

What did the studies actually measure?

Ahrefs, June 2026

Ahrefs looked at 137,210 domains using its web analytics that received traffic in May 2026, checked each root for an llms.txt returning HTTP 200 with markdown rather than an error page, and then read the request logs.

  • 28% of domains (38,360) published a valid file. Ahrefs treats this as an upper bound, since its customers skew technical.
  • 97% of those files were never requested during the month.
  • The 3% that were requested (about 1,100 domains) received all the measured traffic.
  • 96% of requests to the file came from bots, but 77% of that bot traffic came from non-AI tools (monitors, SEO crawlers, scrapers).
  • Among named AI agents, GPTBot led with 4.51% of requests, followed by Claude Code, ClaudeBot at 0.8%, and OAI-SearchBot at 0.74%. PerplexityBot was negligible.

Their conclusion: “If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration.”

SE Ranking, November 2025

SE Ranking measured adoption and then tested for an effect on citations.

  • 10.13% of nearly 300,000 domains had the file.
  • Adoption did not rise with size: 9.88% among sites with under 100 visits, 10.54% in the 1,001–5,000 range, and 8.27% among sites with over 100,000 visits.
  • Their XGBoost model of citation frequency got better when the llms.txt feature was removed: “when we removed the LLMs.txt factor, the model’s predictions actually improved.”

The one-line finding: “There’s no correlation between AI citations and LLMs.txt. Both statistical analysis and machine learning showed no effect of LLMs.txt on how often a domain is cited by LLMs.”

The operators

No major answer engine has committed to reading the file. Google’s Gary Illyes said in July 2025 that Google does not support llms.txt and is not planning to, and John Mueller compared it to the keywords meta tag. OpenAI, Anthropic, and Perplexity document their crawlers in detail and none of that documentation mentions the file.

Why does the file get so little traffic?

Because the retrieval pipelines that produce citations were built on HTML crawling, and there is no step in them that consults a site-level index written for a language model. A search indexer wants every page it can rank. A live fetcher wants the one page the user asked about. Neither has a reason to read a curated table of contents first. The GEO paper summary makes the same point from the other direction: everything that measurably moved citation rate was on the page itself.

Where does llms.txt actually work?

In the one place it was designed for: agents that read documentation. The Ahrefs data shows Claude Code as the second-largest named requester, ahead of Anthropic’s own training crawler. Coding assistants such as Cursor, Continue, and Cline can be pointed at an llms.txt to load clean, current API documentation into a session instead of guessing from training data.

If you publish developer documentation, the file has a real audience and the llms-full.txt companion has an even better one, because a single fetch replaces a crawl. If you publish a restaurant menu, a listing, or a marketing site, nothing reads it today.

Should you still add one?

Yes, if it takes under an hour, for three reasons that have nothing to do with citations:

  1. It is free insurance. The cost is a static markdown file. If any engine starts reading it, you are already there.
  2. Agents are growing faster than search crawlers. The three-kinds-of-crawler split matters here: live fetchers and task agents are the class most likely to consult a site index, and they are the class whose traffic grew most in 2026.
  3. Writing it forces a useful exercise. Naming your ten most important pages in one sentence each is a content audit in disguise, and the result is a good input for the /llms-full.txt file and for your sitemap priorities.

Do not expect a score change on any AI visibility tracker from adding it, and be skeptical of any vendor that promises one.

Why does the audit still weight it at 5 points?

Three reasons, stated plainly so the weight can be judged.

The check was designed in early 2025 when the convention was new and its uptake was an open question. The rubric treats it as one of several agent-readiness signals in the Answer Engine category alongside structured data and readability, not as a proxy for citations. And it remains the only file on your site that coding agents demonstrably request.

The weight is under review in light of the 2026 evidence. If it changes, the change and its reasoning will appear in this knowledge base, and existing audit permalinks will keep the score they were given.

The format, for the record

If you do add one, the format guide covers the spec. The short version: an H1 with the site name, a blockquote summary, H2 sections, linked list items with one-line descriptions, under about 4 KB, served as text/markdown. The audit awards partial credit for a malformed file and full credit for a valid one.


Further reading

See what the audit finds at your root: Run an audit →