Skip to content
AI VisibilityGEO
EN

Most Domains That Win AI Citations Do Not Ship an llms.txt. Most Vendors Selling the Idea Do.

We asked 69 AI visibility vendors and 122 domains we have seen cited for their llms.txt on one day. The vendors ship it far more often than the cited sites.

· Updated · 17 min read

On 14 August 2026 we asked two groups for their llms.txt with the same script on the same day: the 69 vendors in our catalogue of AI visibility tools, and the 122 domains our register has recorded as cited three or more times in AI answers. Of the vendors we could reach, 44 of 68 ship one. Of the cited domains, 49 of 112 do. That was the first arm. It has since been repeated on two more disjoint strata of cited domains, 2,147 in all, and the rate holds three times. The people selling the convention ship it at around two thirds. The sites actually winning the citations ship it at under half.

Disclosure: EchoWi sells AI visibility measurement and ships an llms.txt itself, so we are inside the vendor row and we have an interest in what this file is worth. The method is a GET of /llms.txt on each domain’s own origin, on one day, with challenges and network failures counted separately from absent files. The counts are in our measurement register, the population rule is stated below, and the whole thing takes an afternoon to repeat against us.


The short version

  1. Vendors ship an llms.txt at 44 of 68 reachable domains. That is the group whose product is making you legible to machines.
  2. Domains we have seen cited ship it at 49 of 112. That is the group the machines actually quote.
  3. The gap runs the wrong way for the pitch. If the file were what earned citations you would expect the cited group to be ahead, and it is behind by roughly twenty points.
  4. This is not evidence the file hurts, or that it does nothing. Two populations differing on a convention is not a mechanism, and we are not going to pretend otherwise.
  5. It is enough for one narrow claim, and that claim is useful: winning citations does not require an llms.txt, because most of the domains winning them do not have one.
  6. Six vendor files are under 2 KB, which is present and too small to index anything.
  7. A third stratum of 1,221 once-cited domains settles which half of the crawler cross-tab replicates. Naming does, at z 5.47. Blocking does not, at z 1.87.

What llms.txt claims to be

llms.txt is a proposed convention: a markdown file at the root of a site that gives a language model a curated map of what matters, in the order a human would recommend reading it. The pitch is straightforward and not unreasonable. A crawler landing on a large site has to infer structure from navigation and internal links; a file that says “here are the twelve pages that matter and here is what each one is for” removes that guesswork.

It has no specification behind it in the sense that Content Signals has one, no registry, and no published statement from any major engine that it is read. That last point matters and we will come back to it, because a convention nobody has demonstrated consuming is a different kind of thing from one whose consumer is known.

What it does have is adoption inside one specific industry: the tools that sell AI visibility. That is the observation this piece started from.

Two populations, one day, one instrument, and later a third stratum

The comparison only means anything if both halves were measured the same way at the same time, so that is how it was done.

The vendor population is the 69 tools in our verified catalogue of AI visibility and answer engine optimisation software. It includes us.

The cited population is every domain our measurement register has recorded as appearing in three or more separate probe rows: 122 of the 1,019 distinct domains the register has seen. The threshold is arbitrary and we are declaring it rather than burying it. Three or more means a domain has turned up across several questions rather than once, which is the difference between a site the answer layer reaches for and a site that happened to be there. Lower the bar to two and the population is 286; raise it to five and it is 44.

The instrument is a GET of /llms.txt on each domain’s own origin, following redirects. Three of its choices decide the result:

GET, never HEAD. A server is not obliged to answer a HEAD request at all. An earlier link checker in this project reported two perfectly live pages as dead for exactly that reason, so this asks for the thing it wants to read.

A challenge is not a missing file. Statuses 401, 403, 429 and 999 mean a bot check stood in the way, which is evidence about the edge and not about whether a file exists. Seven cited domains and one vendor land there. Folding them into “no llms.txt” would have manufactured a number in the direction of our own headline, which is exactly when you should be most careful.

A 200 is not a file. Some hosts answer every path with their marketing page. A body that opens an HTML document is not an llms.txt whatever the status line said, and nine cited domains answered that way.

That is why the table below carries two denominators. Asked is the whole population. Reachable is asked minus challenges and network failures. Both are published because a reader is entitled to check whether the choice of denominator is doing the work.

The numbers

VendorsDomains we have seen cited
Asked69122
Reachable68112
Ship an llms.txt4449
Share of reachable65 per cent44 per cent
Under 2 KB63
Median size9,535 bytes15,378 bytes

Two things worth reading off that table before anyone builds an argument on it.

The first is the gap itself: 65 per cent against 44. If shipping this file were the thing that earned citations, the group being cited should be the group that ships it, and the relationship is the other way round.

The second is the median size, which cuts against the simplest dismissal of the file. Among the cited domains that do ship one, the median is over 15 KB: these are not token gestures. And six vendor files are under 2 KB, which is present and too small to be an index of anything, so the vendor number is slightly generous to the vendors.

Asked again on four and a half times the domains, and it holds

The figure above counts cited domains, so that is the axis to widen rather than repeat. The same rule that gave 122 domains on 14 August gives 547 today, because the register has grown, and there is a second stratum of 379 cited exactly twice that shares no member with it. Both were asked on 25 August 2026, with the design and three thresholds written down first.

Cited 3+ timesCited exactly twice
Domains asked547379
Answered whether the file exists503349
Serve one209153
Share of those that answered41.6 per cent43.8 per cent

43.8 is the published figure to the decimal. The first arm comes in at 41.6, and the interval around the original ran from 34.9 to 53.0, so both sit inside it. A number measured on 112 domains held on 503 and again on a disjoint 349.

Crossing it with the crawler rules of the same sites

This is the part the earlier study could not do. The same domains were read for their robots.txt the day before, so publishing an llms.txt and blocking an AI crawler can be crossed on one site rather than compared as two totals.

The prediction, written before the request, was that publishers would block less, because the file exists to serve content to a model. Pooling both strata:

Of the domains that answered bothPublish llms.txtDo not
Domains360479
Name any AI crawler rule34 per cent24 per cent
Disallow at least one AI crawler9 per cent16 per cent

Publishers write AI rules more often and their rules say no less often. Both directions hold pooled, z of 3.37 on naming and 3.14 on blocking. So the file is not decoration: it travels with a posture.

And the arms disagree about which half of that is real, which is worth more than the pooled number. In the frequently-cited stratum the blocking gap is 10 points and clean, z of 3.18, while the naming gap is nothing. In the twice-cited stratum it is the other way round: naming separates hard, z of 4.02, and blocking falls to 3 points, which the frozen design calls distinguishing nothing.

The prediction is therefore confirmed in one arm and not in the other, and that is how it is published. The direction is the same in both. The size is not, and this design cannot say why. A third and larger stratum was run afterwards and it settles this: naming replicates, blocking does not, and the section below has the six cells. Nothing in this paragraph is withdrawn, but do not stop reading here.

And this settles a question a sister study had to leave open. Crossing llms.txt against markdown at vendor level, doing one turns out not to predict doing the other: 62% against 57%, a gap smaller than one vendor. The difference between the two results is not the convention, it is the kind of thing being crossed. Publishing a file and serving markdown are two chores, and their adoption follows cost. What a robots.txt says about an AI crawler is a posture. Crossing two chores separates nothing; crossing a chore against a posture does.

The naming result was not predicted at all. It is reported as an observation, because a finding that arrives after the data is a hypothesis for the next arm and not a result of this one.

A third stratum, and it says which half replicates

The two arms above disagreed, so the honest move was to run the arm with the power to settle it. There are 1,221 registrable domains our register has seen cited exactly once, disjoint from both strata above and larger than the two of them put together. The design, the deciding statistic and four thresholds were frozen and committed before the first request.

Publishers against non-publishersCited 3+Cited twiceCited once
Domains crossed4953441,102
Name an AI rule34 against 2934 against 1635 against 20
z on naming1.244.025.47
Disallow at least one9 against 209 against 117 against 10
z on blocking3.180.891.87

Naming replicates and blocking does not. The frozen rule said that if only one gap survived the largest stratum, that gap is the effect and the other belonged to the stratum it came from. Naming separates in two arms of three, with the two largest z values. Blocking separates in one, the smallest and most heavily cited.

But the direction is the same in all six cells, and that matters more than the thresholds. Publishers name more in every arm, by 5, 18 and 15 points. Publishers block less in every arm, by 11, 2 and 3. What fails to replicate is not the sign, it is the size, and in both rows the heavily-cited stratum is the one that disagrees with the other two.

So the pooled figure above is not withdrawn and it is not enough on its own. A pooled z over strata that disagree by a factor of five is arithmetic, not evidence of a stable effect, and this is the paragraph that says so about our own number.

Pooled over all three strata, naming is 35 against 22 per cent, z of 6.32. We are quoting that one and not its blocking twin, and the reason is the paragraph above rather than the result: naming replicates across the strata, so pooling it means something, and blocking does not, so pooling it would only launder one arm’s number into a bigger denominator.

The prediction that paid was about the publishers

Before measuring we wrote that if publishers were a stable population they would come back near 34 per cent naming and 9 per cent blocking, with bands of 27 to 41 and 4 to 15. They came back at 35 and 7.

Across three disjoint strata spanning one citation to more than three, the publishers give 34, 34 and 35 per cent on naming and 9, 9 and 7 on blocking. That is the same population three times over. It is the non-publishers who move, 29, 16 and 20 on naming, which is why the prediction that they would fall monotonically with citation count failed: there is no trend, the middle arm is simply the odd one.

What the sites are actually writing

The statistic hides something simpler that the tail stratum makes plain. Of the 1,132 domains that served a robots.txt, 306 name an AI crawler and only 104 disallow one. So 204 wrote the rule and said yes. Two block without naming anybody, caught by a blanket User-agent: *.

That reframes the whole cross-tab. The difference between publishers and non-publishers is not that publishers refuse more. It is that they have a written opinion at all, and two times out of three that opinion is a permission.

And 29 of the 104 blockers are a managed default, a block written by a CDN rather than by the site. We separate those because a declaration is not an author, and a count of the web turning off AI training is partly a count of one vendor’s product decision.

The llms.txt rate itself did not move: 41.6, 43.8 and 44.0 per cent across the three strata. Publishing the file is not a property of how much you are cited.

What this does not show

This is the section that matters most, because the headline is the kind that gets quoted without it.

It does not show that llms.txt does nothing. Two populations differing on a convention is not a mechanism. We have not run an intervention, there is no before and after, and nothing here isolates the file from everything else about these sites.

The two populations are not alike, though composition turns out not to be the answer. The cited set contains youtube.com, reddit.com, forbes.com and pcmag.com, large publishers with their own publishing standards and legal review, and our first version of this section said some unknown share of the gap was that. It is not. Splitting the cited population by how often we have seen it cited, the most cited tier ships the file at 8 of 17 reachable and everything below it at 41 of 95, which is the same rate and both far under the vendors. The gap is not the big names.

Selection runs through the middle of it. The cited population is defined by our own probes, in our own markets, on our own questions. A different question set would produce a different 122.

A convention with no demonstrated consumer is hard to test. No major engine has published that it reads llms.txt. We have not observed a fetch of ours by a named AI crawler that we could attribute to the file rather than to ordinary crawling, and we are not going to claim one.

We have not measured whether shipping it changed anything for anyone. That would need the same sites before and after, with a control, which is a different study and a much slower one.

What would actually answer the question

Since we are declining to answer it, it is only fair to say what would.

The clean version is an intervention: take a set of comparable sites that do not ship the file, add it to half of them, hold the other half fixed, and watch citation rates on a frozen prompt set over a period long enough to outlast the noise this register has already measured in AI answers. That noise is large, so the period is long and the sample is not small.

The cheap version, which nobody appears to have published, is server-side: look at your own logs and see whether any named AI crawler requests /llms.txt at all, and at what rate compared with /robots.txt. That is one query against access logs and it would settle the consumer question for a single site. We would rather see ten sites publish that than see another adoption count, our own included.

Common Questions About llms.txt

What is an llms.txt file?

llms.txt is a proposed convention: a markdown file at the root of a domain that gives a language model a curated guide to the site, listing the pages that matter and what each is for. It is not a specification with a registry behind it, and no major engine has published a statement that it reads the file.

How many AI visibility vendors ship an llms.txt?

44 of the 68 vendors we could reach, out of 69 asked, measured on 14 August 2026. One returned a bot challenge, which we count separately from a missing file because a challenge is evidence about the edge rather than about whether the file exists. Six of the files that exist are under 2 KB, which is too small to index a site of any size.

Do the sites that get cited in AI answers ship an llms.txt?

Less often than the vendors do. Of 122 domains our register has recorded as cited three or more times, 112 were reachable and 49 ship one, which is 44 per cent against the vendors’ 65. That is a comparison of two populations rather than a mechanism, and the two are not alike, so it does not show the file is useless. It does show that winning citations does not require one.

Does llms.txt help with AI search visibility?

Nobody has published evidence either way that we can find, and this study does not settle it. What it rules out is the strongest version of the claim: the file cannot be necessary for citation, because most of the domains that win citations do not have one. Whether it helps at the margin would need an intervention with a control, not an adoption count.

Should I add an llms.txt to my site?

It is cheap, it is not harmful, and if your site is large and badly structured it is a reasonable thing to hand a crawler. What we would not do is buy it as a visibility product, or let it displace work on the things this register has measured actually varying with citation, which are what a page says and whether anyone else says it too. Check your own access logs first: if no AI crawler has ever asked for the file, you have learned more in one query than any adoption count can tell you.

How do I check what a domain ships?

Request /llms.txt from the domain’s own origin with a GET rather than a HEAD, and read the body before believing the status code. If it opens with an HTML document you have been handed a marketing page, not a file. If the response is 401, 403, 429 or 999 you have hit a bot challenge and learned nothing about whether a file is there.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)