A census of robots.txt or crawl-preference signals counts strings in files, but a string does not record who wrote it. Increasingly the most common signal values are not an operator's considered choice; they are a content delivery network's managed default, written into a large number of files at once from a single setting. This note takes a case already measured in this publication's corpus — 33 of the 42 Content-Signal lines found carried one CDN's byte-identical managed default — and draws the measurement lesson: a reported adoption rate for any crawl-preference signal conflates operator intent with platform default unless the two are separated. It reads the 2025 shift toward default blocking and pay-per-crawl, including the revival of HTTP 402 as an enforced, request-time refusal distinct from advisory robots.txt, as the reason the confusion is growing, and proposes disclosing a signal's authorship as a measurement requirement.
A census that measures crawl policy reads files. It counts the domains whose robots.txt disallows a named crawler, or carries a preference signal, and reports the rate. The unstated assumption is that each file records a decision its operator made. That assumption is weakening. A content delivery network sitting in front of a large fraction of the web can write a default robots.txt, or a default preference signal, into every site it serves — from one setting, applied to domains that never touched the file themselves. When it does, a census measuring “adoption” is partly measuring one vendor’s default, not the aggregate of operator choices it appears to report. The first question to ask of any crawl-policy rate is therefore not how high it is, but who authored the signals it counts.
The evidence already in hand
This is not hypothetical for the corpus behind this publication. The companion note on Content Signals in robots.txt reports a re-analysis of the BA-D-2026-01 crawler-access census: of the 42 domains carrying a Content-Signal line, 33 — 79% — set the byte-identical value search=yes, ai-train=no, use=reference, the managed default of a single CDN, which that vendor documents applying by default with search=yes and ai-train=no [4]. The nine hand-authored files, by contrast, varied among themselves and stayed within the policy’s three documented keys. The single most common crawl-preference string in that sample is thus not what operators think about the policy; it is one platform’s default, reproduced. A rate reported without that split — “3.0% of parsed domains express a content signal” — is arithmetically correct and interpretively misleading, because four of every five carriers expressed nothing they authored.
Why the hazard is growing
The mechanism that produces such defaults has moved to the center of the CDN’s product, which is why a measurement caution now applies to a larger share of the web than it once did. On July 1, 2025, Cloudflare announced it was “changing the default to block AI crawlers unless they pay creators for their content” [2], and that “upon sign-up with Cloudflare, every new domain will now be asked if they want to allow AI crawlers, giving customers the choice upfront to explicitly allow or deny AI crawlers access” [3]. The choice is offered, but a default still stands behind it: a domain that accepts the onboarding default ships a crawl policy it did not compose, and the platform’s managed robots.txt and managed content-signal features write the file’s bytes directly. The observation here is purely mechanical, and no evaluation of the policy is intended: one actor can now author the crawl-facing policy of a very large number of sites at once, and a census cannot tell, from the bytes alone, which of those sites chose the words.
A second surface: the metered response
Part of the same 2025 shift changes what a census must look at, not only how it reads what it finds. Alongside default blocking, Cloudflare introduced “pay per crawl,” which it describes as dusting off “HTTP response code 402”: in the design, an AI crawler requesting a paid URL either “present[s] payment intent via request headers for successful access (HTTP response code 200), or receive[s] a 402 Payment Required response with pricing” [1]. Whatever becomes of the commercial arrangement, the measurement consequence is that crawl control is moving partly out of robots.txt — a file a census fetches once, from any client — and into the response to the crawl itself, a status code visible only to a client that makes the request as the metered crawler would. A study that counts robots.txt disallow rules and stops there will not see a site that serves its content to a browser and returns 402 to a crawler. Crawl policy now has at least two surfaces, and they differ in kind: the advisory file, whose rules RFC 9309 states “are not a form of access authorization” [5], and the enforced, request-time refusal a 402 represents. “Who is blocked” is no longer answerable from the file alone.
The measurement lesson
The lesson generalizes past any one vendor. Whenever a platform can inject a default signal, an adoption rate measured over files becomes a mixture of two populations — operators who authored a choice and operators who inherited a default — and the mixture proportion is itself a finding, often the more important one. Three disclosures follow for anyone reporting a crawl-policy rate. First, report authorship: separate hand-authored signals from platform-managed defaults, using whatever fingerprint the managed text leaves — here, a value byte-identical across carriers and an extra key the announced policy never defined. Second, report the response axis, not only the advisory file: whether the measurement observed status codes such as 402, and as which client it fetched. Third, treat a signal’s provenance as part of its meaning: search=yes, ai-train=no authored by an operator and the same string shipped as a default are different measurements that happen to share a spelling. These extend, to the authorship of a signal, the disclosure discipline this publication already requires of any visibility claim in minimum disclosure requirements (BA-C-4) — a rate is not evaluable until the reader knows what was actually counted.
Limitations
This note draws its one worked example from a single corpus (the BA-D-2026-01 census, four sectors, the 1,381 domains that parsed) and a single CDN’s default, identified by a byte-identical string; it does not estimate how common managed defaults are across the web, and other platforms leave different fingerprints or none. A more comprehensive signals census now in preparation will widen this sample, but does not change the caution. The vendor announcements cited are dated July 1, 2025 and were read on 2026-07-10; defaults, onboarding flows, and the status of the pay-per-crawl beta all change, so “the default is X” is a perishable finding, not a fixed property. The note describes a measurement hazard and the disclosures that address it; it evaluates no vendor’s policy, pricing, or the legal effect of a payment-gated response, and takes no position on whether default blocking or pay-per-crawl is desirable. Finally, separating authored from inherited signals depends on the managed text being recognizable: where a platform default is indistinguishable from a hand-authored file, the split this note asks for cannot be made, and that limit should itself be disclosed.
References
- 1.Cloudflare (W. Allen, S. Newton). Introducing pay per crawl: enabling content owners to charge AI crawlers for access (2025). https://blog.cloudflare.com/introducing-pay-per-crawl/ Accessed 2026-07-10. [archived]
- 2.Cloudflare (M. Prince). Content Independence Day: no AI crawl without compensation! (2025). https://blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/ Accessed 2026-07-10. [archived]
- 3.Cloudflare (press release). Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for a New Business Model (2025). https://www.cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/ Accessed 2026-07-10. [archived]
- 4.Cloudflare (W. Allen). Giving users choice with Cloudflare's new Content Signals Policy (2025). https://blog.cloudflare.com/content-signals-policy/ Accessed 2026-07-10. [archived]
- 5.M. Koster, G. Illyes, H. Zeller, and L. Sassman, IETF. RFC 9309: Robots Exclusion Protocol (2022). https://www.rfc-editor.org/rfc/rfc9309.html Accessed 2026-07-10. [archived]
How to cite
PDF of recordBarkhausen AI (2026). When the CDN writes your policy. https://barkhausen.ai/notes/cdn-managed-crawl-policy/
BibTeX
@techreport{BA-MN-6,
author = {{Barkhausen AI}},
title = {When the CDN writes your policy},
institution = {Barkhausen AI},
year = {2026},
url = {https://barkhausen.ai/notes/cdn-managed-crawl-policy/}
}Published under the Creative Commons Attribution 4.0 International (CC-BY-4.0).
