Skip to content
AI VisibilityGEO
EN

48 of the 76 Sites Refusing AI Training Did Not Write That Refusal

Across three strata of cited domains, 48 of the 76 refusals of AI training were a CDN default, and 53 of the 81 signals a person typed permit it.

· Updated · 26 min read

A Content Signal is the one machine-readable line that says what an AI system may do with your page after it reads it. We asked three disjoint strata of cited domains for theirs. 2,002 served a robots.txt and 129 declare a signal, and the interesting number is neither: of the 76 that refuse AI training, 48 did not write that refusal, and of the 81 declarations a person did type, 53 permit training.

Every count of “the web is turning off AI training” that is read from robots.txt at scale is, in part, counting a vendor’s product decision. Here is how large that part is on one measurable population.

Disclosure: EchoWi sells AI visibility measurement, so a study about who declares their preferences properly is a study with an interest. The method is one GET request for /robots.txt per domain and a check of whether each Content-Signal: line falls inside a # BEGIN Cloudflare Managed content block. The date is 24 and 25 August 2026, all three populations and every threshold were written down before their first request, every declaring domain is named in our measurement register, and anyone can repeat it against us in a minute.


The short version

  1. 547 domains asked, 511 served a robots.txt. 26 answered a bot challenge, 5 could not be reached, 5 had no file. Those three are counted apart, because a challenge is not an absent file.

  2. 31 of the 511 declare a Content Signal. That is 1 in 16, on a population selected for being cited by AI answers.

  3. 10 of the 31 carry every line inside a Cloudflare managed block. Nobody at those sites typed them.

  4. All 10 managed lines are the same string, byte for byte. One default, not ten decisions.

  5. 16 domains refuse AI training. For 10 of them, the refusal is the default. That is the finding: declaring is mostly a human act, refusing mostly is not.

  6. And among the 21 that did type something, there are only 6 distinct wordings. Three of them cover 18 domains. That part was not predicted in the first arm and is reported as an observation there.

  7. Repeated on 379 disjoint domains, it holds. 24 declare, 9 are the managed default, the managed wording is the same one string, and the wording tally, frozen as a prediction this time, came back at 3 wordings covering 11 of 15.

  8. Asked of a second convention, the same twenty sites turn up. Of the 120 domains that disallow a named AI crawler, 20 do so only because of the managed block, and 19 of those 20 are the same domains whose Content Signal was the default. One product feature writes both.

  9. The default is not a switch, it is a position. It shuts out GPTBot, ClaudeBot and Google-Extended while leaving OAI-SearchBot, Claude-User and Googlebot alone: block the training, keep the search, on twenty sites that did not write it.

  10. And on a third stratum of 1,221 once-cited domains, the mechanism turns over. Pooled across all three, 48 of the 76 sites refusing AI training did not write that refusal, and of the 81 declarations a person did type, 53 permit AI training. Two of every three human-typed signals say yes; the machine-written one always says no.


Why authorship is the question, and not adoption

The convention itself is simple. A line in robots.txt says search=yes, or ai-train=no, or ai-input=yes, and a machine reading the file learns what the publisher permits. We asked the AI visibility vendors the same question ten days earlier, and the answer there was 9 of 80.

What changed between the two studies is not the question. It is that the instrument learned to ask a second one.

Cloudflare, which proposed the convention, can insert a managed block into a customer’s robots.txt. That block carries its own comment header and its own Content-Signal line. Read flat, a declaration inside it looks exactly like a declaration a publisher typed. The file says the same thing either way, and the count is the same either way.

The two are not the same claim. “This publisher decided not to allow training” and “this publisher’s CDN has a default, and the publisher has not changed it” describe different worlds, and only one of them is evidence about what publishers want. Separating them costs nothing once you look for the block markers, and it changes the headline.

What we asked

The population is the registrable domains this register has recorded as cited in three or more places across its measurements, collapsing www. and subdomains to the registrable name. That is 547 domains. All of them were asked. None were selected.

Asked547
Served a robots.txt511
Answered a bot challenge26
Could not be reached5
Served no file5

The three buckets after the first are kept separate on purpose. A challenge is a fact about the edge, not about the file, and a host that never answers is a fact about our reading rather than about the site. Folding either into “does not declare” would inflate the silent majority with sites we simply did not read.

31 of the 511 declare at least one Content-Signal line. One in sixteen, and that is on a population chosen for having been cited in AI answers, which if anything skews toward sites that pay attention to this.

A third of the declarations were typed by nobody

The prediction, frozen before the first request, was that 60 per cent or more of declaring domains would carry every line inside a managed block, with 30 per cent or less retracting the idea entirely.

The answer is 10 of 31, or 32 per cent, which lands one point above the retraction threshold and is published as what it is: partial, and much smaller than predicted. Most declarations on this population are not a CDN default. About a third are.

That is worth stating plainly because the opposite headline was available and would have been more dramatic. It is not what the data says.

The second prediction held exactly. All 10 managed declarations carry the identical string, search=yes,ai-train=no,use=reference, spacing included, and in every file we opened it sits on the same line of a structurally identical block. It is one default, deployed 10 times, not 10 sites converging on the same wording.

The refusals are the other way round

The third prediction is where the study earns its title.

Of the 31 declaring domains, 16 refuse AI training: their line contains ai-train=no. Split by who wrote it:

Refusals typed outside a managed block6
Refusals that are the managed default10

So the two statistics point in opposite directions on the same population. Declaring is mostly a human act, and refusing mostly is not. Two thirds of the refusals on this sample were placed by a default that the site owner may never have seen, let alone chosen.

Nothing here says those publishers disagree with the refusal. A default can express someone’s wishes perfectly well. What it cannot do is serve as evidence that they decided anything, and a headline of the form “N per cent of sites now block AI training” is exactly that kind of evidence claim.

Six wordings for twenty-one declarations

This next part was not predicted, and it is reported as an observation rather than a result, because the analysis was chosen after seeing the data.

Among the 21 domains that declare outside a managed block, there are 6 distinct wordings, and 3 of them account for 18 domains. Seven unrelated companies publish ai-train=yes, search=yes, ai-input=yes. Six publish the same three fields in a different order. Five publish ai-train=no, search=yes, ai-input=yes. Only 3 domains in the whole set carry a wording that appears once.

We checked whether the seven-domain string is one content management system emitting one file, because that would have been the cheap explanation. It is not. Their robots.txt files run from 9 to 130 lines and the Content-Signal line lands anywhere from the first line to the 132nd, in completely different contexts.

Then we checked the obvious source, which cost one request. contentsignals.org, the reference site for the convention, publishes that exact string with that exact spacing. So the likeliest reading is that seven publishers copied the example, which is a normal and reasonable thing to do and is also not the same as writing down a preference.

Run again on a sample it was not found in

A figure found in one sample is a figure about that sample until it turns up in another. So the same questions went to a disjoint stratum: the 379 registrable domains this register has seen cited exactly twice, which share no member with the 547 above. The four predictions and their thresholds were written down before the first request.

361 served a robots.txt, 12 answered a challenge, 5 could not be reached, 1 had no file. 24 declare a Content Signal.

Arm 1, cited 3+ timesArm 2, cited exactly twice
Domains asked547379
Declaring3124
Declarers that are all managed10, or 32 per cent9, or 38 per cent
Distinct managed wordings11
Refusing AI training1619
Refusals that are the default109, or 47 per cent

The managed share replicates. The prediction was 17 to 47 per cent, meaning within fifteen points of the first arm, and it came back at 38.

The single wording replicates. All 9 managed declarations in the second arm carry the same string as the 10 in the first.

The refusal claim replicates in direction and not at the level predicted, and that is worth saying exactly. The prediction as written was “half or more”, and 9 of 19 is 47 per cent, so it misses by one domain. The retraction threshold was one in five, and it is nowhere near that. So the honest sentence is that the default accounts for close to half the refusals in both arms, not that it accounts for most of them.

And the wording tally, which was an observation yesterday, was frozen as a prediction and held. In the second arm, 15 domains declare outside a managed block using 7 wordings, and the 3 commonest cover 11 of them, or 73 per cent, against a threshold of half. Pooling only that claim, because it is the one both arms make about the same object: 36 typed declarations, 10 distinct wordings, 3 of which cover 28.

The most common typed wording across both arms is not the permissive one. It is ai-train=no, search=yes, ai-input=yes, on 11 unrelated domains, byte for byte.

The same question of a second file, and the same twenty sites

A Content Signal is one line of a young convention. The old way to tell an AI crawler to stay out is a Disallow rule with the crawler named, and it has been available for years. If authorship is a real property and not an artefact of one convention, it should show up there too.

So the same 870 served files were read again, this time for the verdict that fourteen named AI crawlers get at the site root, and read a second time with the managed block removed, to see which verdicts the site’s own lines produce.

Arm 1Arm 2
Domains serving robots.txt510360
Disallowing at least one AI crawler8238
Of those, the block is only the managed text1010
Domains naming no AI crawler at all349271

The prediction, written before the request, was that at most one blocker in five would be the default, on the reasoning that a named Disallow is an older and more deliberate act than a signal line. It came back at 12 per cent in the first arm, comfortably inside.

The number was right and the reasoning was wrong, which is worse than being wrong, because nobody rechecks it.

The twenty domains whose crawler block is only the managed text are, to nineteen, the same domains whose Content Signal is only the managed text. Not a similar population. The same list. So the difference between 12 per cent of blockers and 32 per cent of declarers is not a fact about Disallow rules being more deliberate than signal lines. It is the same twenty sites divided by two different denominators.

What the default actually says

Here is the block, from one of the twenty, with nothing removed:

# BEGIN Cloudflare Managed content

User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /

User-agent: Amazonbot
Disallow: /
...
User-agent: Google-Extended
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

# END Cloudflare Managed Content

It allows everything to everyone, then names nine crawlers and shuts each out. Read the list and it is a position, not a switch: GPTBot is out and OAI-SearchBot is not. ClaudeBot is out and Claude-User is not. Google-Extended is out and Googlebot is not.

That is block the training, keep the search, expressed identically on twenty sites in two disjoint strata, and none of the twenty wrote it.

Whether those publishers agree with it is a separate question this cannot answer, and a default can reflect a site owner’s wishes perfectly well. What our instrument sees is who typed the line, and that is all we will claim from it.

Why it matters that the default has an opinion

Every reading of these files that treats them as expressed preference now has a floor under it. Not because publishers are careless, but because the layer is small enough that one product feature moves it.

There is a second-order effect worth naming, and it points the other way from the usual worry. This register measures, on the same population, that training crawlers are refused far more often than search crawlers. The managed default is itself a block-training-keep-search policy, so part of that gap could be the default rather than the web. Removing the managed text and re-counting, the gap narrows and holds: GPTBot goes from 3.4 times OAI-SearchBot to 2.6, and ClaudeBot from 2.5 times Claude-User to 1.9.

The direction is the web’s. Some of the size was the vendor’s.

The confounder the design named in advance

A domain cited twice is a less-cited stratum than one cited three or more times, and the two arms differ on something the design did not predict: 19 of 24 declarers refuse training in the second arm, against 16 of 31 in the first. That is 79 per cent against 52.

This design cannot say whether that is the stratum, the kind of site, or chance on small numbers. It is written here rather than explained, because a difference this size deserves naming even when the study that found it cannot resolve it.

What “typed” can and cannot mean

This is the limit that matters, and the wording tally above is what makes it measurable rather than rhetorical.

Our instrument separates lines inside a Cloudflare managed block from lines outside one. That is a real distinction and it is exactly what the block markers let us see. It is narrower than authorship. A line pasted from a reference site, a line added by a plugin, a line inherited from a template are all counted as typed, because none of them sits in a managed block.

So 21 is a ceiling on the number of authored declarations in this sample, not an estimate of it. The floor, if you treat every repeated wording as copied, is 3. The truth is somewhere between, and no reading of robots.txt alone can locate it.

We would rather publish that range than the ceiling with a confident label on it.

What the root-only verdict costs, measured

The limit below says a crawler verdict here is the verdict at the site root, and a domain could always disallow a crawler on a section instead. That is worth a number rather than a caveat, and the number came from the files the sweep already had, so it cost no request.

Every one of the 120 blockers uses a bare Disallow: /. Not one narrows the block to a section. Two routes agreed on that, the parser and a hand-rolled scan of the raw groups, which is how this register checks a claim it intends to lead with.

The other half is the domains that let every AI crawler through at the root. There are 750 of them, and ten common paths were put to each: article and blog sections, a search page, an API prefix and an uploaded file.

Of the 750 that allow every AI crawler at the root
Disallow one on a probed path130
…where the same rule also stops Googlebot114
…where Googlebot is allowed and the AI crawler is not16

That middle row is the whole point. A path rule that also stops Googlebot is an ordinary robots rule about a section, not a policy about AI, and counting the two together reports 130 AI path blocks where there are 16. A sweep that did not ask whether the rule is AI-specific would overstate this layer eightfold.

The 16 are /api/, /search and /wp-content/uploads/: infrastructure, not editorial. So the root verdict is the whole answer for every blocker, and it misses about one in fifty of the rest, on paths nobody publishes articles under.

Ten paths is ten paths, so 16 is a floor and not a count.

A third stratum, and the blocking rate falls with citation depth

The two arms above were 547 and 379 domains. On 25 August we read the same two files for the 1,221 registrable domains our register has seen cited exactly once, a stratum disjoint from both and larger than the two together. Same fourteen agents, same managed-block detector, same day as its llms.txt twin.

Of the domains that served a robots.txtCited 3+Cited twiceCited once
Served5103601,132
Disallow one of the published ten15.5 per cent10.3 per cent9.0 per cent
Name any AI crawler rule32 per cent25 per cent27 per cent
Of the blockers, managed default12 per cent26 per cent28 per cent

The blocking rate falls monotonically with how often we have seen a domain cited, and the naming rate does not. Three disjoint strata, one instrument, one day. That is the opposite of what a “blocking costs you citations” story predicts, and it is not a causal claim: the obvious reading is that the domains we see cited most often are large publishers, and large publishers are the ones with a legal position on AI training.

And the share of blocking that nobody typed runs the other way. Among the sites that block, the proportion whose block is entirely the managed default goes 12, 26 and 28 per cent as the stratum gets less cited. So the deeper into the tail you look, the more of the web’s refusal is a CDN’s product decision rather than a publisher’s decision, which is the finding of this article measured on a population it was not found in.

Against the thresholds this study froze before its first request: the published-ten rate lands at 9.0 per cent, below the 11.6 lower bound of the interval around the originally published 20 per cent, so that figure does not describe this stratum and the article says so rather than averaging it away. The Googlebot control passes at 1,130 of 1,132, and the most blocked agent is a training crawler again, CCBot at 80.

The tail stratum, and two of every three typed signals say yes

The published study read two strata. A third was frozen before the first request, on the 1,221 registrable domains this register has cited exactly once, disjoint from both arms above, and it was chosen because the crawler half of this piece had just measured a gradient: the managed share of blockers rises as citation falls. If that shape belongs to the web rather than to one convention, it should appear here too.

PredictedThresholdMeasured
Declaration rate4 to 9 per centoutside the band6.5 per cent, 74 of 1,130
Managed share of declarers40 per cent or more30 or less kills it39 per cent, partial
Managed share of refusers45 per cent or more25 or less retracts the headline71 per cent, 29 of 41
Top three typed wordingshalf or more73 per cent, 33 of 45

Adoption replicates three times and almost to the decimal. 6.1 per cent, 6.6 and 6.5, across three disjoint strata running from one citation to more than three. Writing a Content Signal is not a property of how often you get cited.

The second prediction landed one tenth of a point short of its threshold, and is published as partial. The bar was 40 and the figure is 39.2. The direction is monotone, 32, 38 and 39, and the strength does not arrive, so the gradient cannot be called cross-convention. Rounding 39.2 to “basically 40, confirmed” would have been choosing the reading after seeing the data.

The fourth held at exactly the 73 per cent of the second arm, which is twice now that a prediction promoted from an after-the-fact observation has come back on a sample it was not found in.

The finding the design did not predict

Pool the three disjoint strata and the study says something it could not say with one: of 2,002 domains serving a robots.txt, 129 declare a Content Signal, 76 of those refuse AI training, and 48 of the 76 did not write that refusal.

Now turn it over. 81 of the declarations were typed by a person, and 53 of those 81 permit AI training.

Two of every three signals a human types say yes. The one a machine writes always says no. The managed default emits a single string and that string is a refusal, so every managed declaration refuses: it saturates the refusers and dilutes among the declarers. That is exactly why the second prediction came in partial while the third came in high, and the frozen design said in advance that such a split would mean the difference is in what is declared rather than who declares it.

So a count of “the web is switching AI training off,” read from robots.txt at this scale, is measuring one vendor’s product decision more than it is measuring an opinion. The people who bother to type a preference mostly say come in.

What the instrument did, said out loud

The sweep was run three times the same day, because it costs nothing to repeat. Declarers 74, managed 29, typed 45, refusals 41 and managed wordings 1 on all three. The only figure that moved is the denominator, 1,131, 1,132 and 1,130, as one or two hosts flickered between answering and not. A statistic whose numerator repeats three times and whose denominator wobbles by two is worth much more than a single reading.

The refusal count is a floor by exactly one, and the detector was left alone on purpose. One site writes indexed-content=allow, ai-training=disallow, ai-search=allow, which are not the keys Content Signals defines: the specification knows search, ai-input and ai-train. Our detector looks for ai-train=no, so that row is not counted as a refusal even though the intent is plain. Widening the detector would move the two published arms without anyone knowing how many of their rows use another vocabulary, so the exception is declared instead.

And one more thing only a person writes: a site publishes ai-train=si, the value in Spanish inside a specification that knows only yes and no. A machine does not make that typo. It is the cheapest evidence there is that the 45 typed declarations were typed by somebody.

What this does not show

Two conventions, two days, three populations. 547 domains cited three or more times, 379 cited exactly twice and 1,221 cited exactly once, read on 24 and 25 August 2026 for their Content Signal and for their crawler rules, and all three selected by citation, so none is a sample of the web and no share here transfers to one.

Declaring is not complying, and not declaring is not blocking. A Content Signal states a preference. It does not enforce anything, a crawler is free to disregard it, and the 1,873 domains that declare nothing have not thereby denied anything. Silence is silence.

A managed default is not a disagreement. We can see who typed the line. We cannot see whether the publisher agrees with it, whether they enabled the feature deliberately, or whether they have ever read their own robots.txt.

A robots.txt is what the edge serves, not what the repository holds. We know this one from our own experience: a declaration of ours was once invisible to a sweep because a CDN was serving a version cached with a year of TTL.

The two arms differ on something neither predicted. The less-cited stratum refuses training far more often, 19 of 24 declarers against 16 of 31, and this design cannot say whether that is the stratum, the kind of site, or chance on small numbers.

A crawler verdict here is the verdict at the site root. A domain that allows GPTBot at / may still disallow it on a section, and a domain counted as blocking may allow a path we did not ask for. The root is the comparable question across 870 files; it is not the whole file. What that costs is measured above rather than left as a caveat.

And we have not measured whether any of this changes citation. Nothing here links a declaration to being cited more or less often. That needs a different design, and saying so is cheaper than implying the link.

What a publisher should take from it

If you have a Content Signal you did not write, you have a stated position on AI training that you did not choose. That is worth thirty seconds: open your own /robots.txt, search for Content-Signal, and see whether the line sits inside a block whose comment names your CDN.

If it does, the decision is still available to you. Turn it into your own line, or leave the default in place deliberately. Either is fine. What is not fine is finding out from somebody else’s study.

And if you are reading a headline about how much of the web has turned off AI training, the question to ask is the one this study exists to ask: how many of them typed it?

The same discipline is what separates a visibility number you can act on from one you cannot. Knowing that an assistant named you is not knowing that it cited your page, and knowing that a preference is stated is not knowing that a person stated it. That distinction is what our platform measures, one surface and one market at a time, and it is the reason this study could be run at all.

Every figure here sits alongside every other measurement we publish, each with its date, its population and its method, in our measurement register.

Common Questions About Content Signal Authorship

What is a Content Signal?

A line in robots.txt that states what automated systems may do with a page after they fetch it. The vocabulary comes from a proposed IETF standard drafted by the AI Preferences working group, and the fields in common use are search, ai-input and ai-train, each set to yes or no. It expresses a preference and enforces nothing.

How many websites declare one?

31 of the 511 domains we could read, on 24 August 2026. That is roughly 1 in 16, on a population of 547 registrable domains selected for having been cited in AI answers. It is not a sample of the web, and a share measured on a different population would be a different number.

Can a CDN add a Content Signal to my robots.txt?

Yes, and 10 of the 31 declarations we found are exactly that. Cloudflare can insert a managed block into the file it serves, marked with a # BEGIN Cloudflare Managed content comment, and the Content-Signal line inside it is a platform default rather than something the site owner wrote. All 10 in this sample carry the identical string.

How do I tell whether I wrote my own Content Signal?

Open your own /robots.txt and look at what surrounds the line. If it sits between comments naming a managed block, your CDN placed it. If it sits among your own rules, somebody at your site put it there, though that still does not tell you whether they wrote it or pasted it from a reference page.

Does refusing AI training in robots.txt actually stop training?

No, and nothing in this study measures compliance. A Content Signal is a stated preference in a file that crawlers may read and may ignore. What it does give you is a public, dated, machine-readable record of what you asked for, which is a different and more modest thing than enforcement.

Why does it matter who typed the line?

Because a count of stated preferences is being read as evidence of publisher intent. On this population, declaring is mostly a human act and refusing mostly is not: 10 of the 16 refusals are the same default string. Any conclusion of the form “publishers are turning against AI training” needs to subtract the ones nobody chose.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)