On 2026-07-09 the robots.txt bytes behind this publication's crawler-access corpus (BA-D-2026-01) were re-scanned — no new collection — for four emerging AI-usage preference signals: a Content-Signal line, an x-rsl field, a content-usage field, and a license field. The scan is a raw-text line scan over the 1,381 domains whose robots.txt parsed cleanly. The signals are sparse and sector-concentrated: Content-Signal appears on 42 domains (3.0%), x-rsl on 7 (all news), content-usage on 3 (all e-commerce), license on 3 (all news). The organizing finding is authorship: of the 42 Content-Signal carriers, 33 (78.6%) carry one CDN's byte-identical managed default and 9 are hand-authored, so most deployment is a platform default rather than an operator-authored string. This completes the single-signal preview in the Content Signals note. It measures deployment, not compliance, and reports CDN penetration as a lower bound.
1. Summary
On 2026-07-09, the raw robots.txt bytes behind this publication’s crawler-access census (BA-D-2026-01) were re-scanned for four emerging AI-usage preference signals — a Content-Signal line, an x-rsl field, a content-usage field, and a license field. There was no new collection: the analysis reads the identical stored robots.txt responses that BA-D-2026-01 fetched and evaluated, so this census is a derived measurement of the same corpus on the same day. The scan is a raw-text line scan — it counts the presence and literal value of each field, not a conformance parse — over the 1,381 domains of that corpus whose robots.txt parsed cleanly (universities 398, news 429, e-commerce 303, government 251). The unit of analysis is the presence of each signal in one file, as served on one day.
The census has one organizing finding, and it is about authorship rather than adoption. A Content-Signal line — the most common of the four emerging signals — appears on 42 of the 1,381 parsed domains (3.0%). But 33 of those 42 carriers (78.6%) carry one content-delivery network’s managed default: a managed-robots.txt comment signature together with the byte-identical value search=yes, ai-train=no, use=reference. Only 9 are hand-authored. The most common Content-Signal string in the corpus is therefore not 42 separate expressions of a preference; it is, in 33 cases, a single platform’s default text reproduced unchanged, and in 9 cases an operator-authored line. Separating the two is the point of this census.
This measurement was previewed. The Content Signals note reported the 3.0% figure and the 33-of-42 split as a first, single-signal field measurement drawn from this same corpus; this census is the full signal-layer measurement those numbers were a preview of, and it publishes the per-domain dataset. The measurement lesson — that any adoption rate for a crawl-preference signal conflates operator intent with platform default unless the two are separated — is developed in the method note When the CDN writes your policy (BA-MN-6). This census reports the measured facts: the topography of the signal layer, its authorship decomposition, and the CDN-penetration floor that bounds both. Two disciplines run through it. The denominator is stated everywhere and is the parsed corpus of 1,381, not the full 2,000 (Section 2.1). And every figure is a count of what a file contains, never a claim that any crawler honors it — deployment, not compliance.
2. Scope and method
2.1 The corpus and the parsed denominator
The corpus is the four-frame, 2,000-domain set of BA-D-2026-01: top-500 selections of universities, news outlets, e-commerce sites, and U.S. federal government domains, each built from a cited public source under a documented ordering, and every frame fully enumerated rather than sampled. The frame construction, the fetch, and the category assignment are specified in that census and are not restated here; the crawler-access dataset carries the row-level fetch outcomes. Because this analysis reuses BA-D-2026-01’s stored bytes unchanged, its frames, its fetch window, and its non-response are exactly those of the parent census.
The denominator throughout is the parsed subset: the 1,381 domains whose stored robots.txt was a real, readable file (BA-D-2026-01’s parsed category). This is a narrower denominator than the parent census’s access figures, which are computed over knowable domains (parsed plus no_robots_file). The reason is mechanical: an emerging signal is a line inside a robots.txt file, so a domain with no file (no_robots_file) has nowhere to carry one, and a domain whose file could not be read (unparseable_blocked, fetch_error) has an unreadable file rather than an absent signal. Restricting to parsed files makes “0” mean “a readable file that does not contain this line,” not “no file.” The consequence is a selection this census states plainly and returns to in Section 5: the 383 unparseable_blocked domains — many sitting behind a bot-challenge page — are exactly the domains most likely to be heavy CDN users, and their signals are invisible here, which biases the CDN-related counts downward.
2.2 The four signals and the line-scan stance
Four emerging fields were scanned, each recorded as a presence boolean and, for Content-Signal, a literal value:
Content-Signal— the line defined by the Content Signals Policy [1], a comma-separated list of purpose signals (search,ai-input,ai-train) each set toyesorno. The policy and its axis are described in the Content Signals note; this census measures its presence and value in the wild.x-rsl— anx-rslfield observed in the raw bytes, one of several licensing-oriented header conventions.content-usage— acontent-usagefield, a separate usage-preference vocabulary.license— alicensefield naming a content licensing statement.
The x-rsl, content-usage, and license fields are reported as observed raw field names — distinct emerging vocabularies present in the same files — not against any one specification, and no conformance or licensing effect is claimed for them here. All four are measured by a raw-text line scan: the analysis matches each field by line, records whether it is present, and for Content-Signal records its literal and normalized value. It does not parse the file under RFC 9309 [3], does not resolve directive grouping or precedence, and does not judge whether any value is semantically valid. A count is “this line exists in this file,” not “this preference is in force.”
2.3 The authorship classification
Each Content-Signal carrier is classified as CDN managed default or hand-authored by a two-part test, and the two parts matter because either alone would misclassify:
- CDN managed default requires both the CDN’s managed-
robots.txtcomment signature (a known managed-content marker in the file’s comments) and the byte-identical valuesearch=yes, ai-train=no, use=reference— the value that CDN documents as its managed default [2], carrying a fourth key,use, beyond the policy’s announced three [1][2]. 33 of the 42 carriers meet both conditions. - Hand-authored is the complement: a
Content-Signalline without the managed signature, whatever its value. The 9 hand-authored carriers use varying values (Section 4).
The classification is deliberately not “Server header equals the CDN.” As Section 4 shows, 41 of the 42 carriers sit behind that CDN by Server header, including 8 of the 9 hand-authored files, so a Server-header split would misassign nearly every hand-authored file to the managed bucket. The managed-comment signature plus the exact default value is what separates a default that was injected from a line that was written; the Server header alone does not.
3. The signal layer is sparse and sector-concentrated
The first result is how little of this signal layer exists in the wild, and how unevenly it is distributed. The table reports each emerging signal’s presence count over each sector’s parsed domains, with the parsed denominators in the header. Two established directives — crawl-delay and a sitemap line — are shown beneath a rule as context, because they set the scale against which the emerging signals are sparse; they are not part of the emerging-signal story and are treated in their own notes (crawl-delay).
| Signal | Universities (n=398) | News (n=429) | E-commerce (n=303) | Government (n=251) | Total (n=1381) |
|---|---|---|---|---|---|
| Content-Signal | 10 | 16 | 8 | 8 | 42 |
| x-rsl | 0 | 7 | 0 | 0 | 7 |
| content-usage | 0 | 0 | 3 | 0 | 3 |
| license | 0 | 3 | 0 | 0 | 3 |
| crawl-delay (context) | 54 | 61 | 69 | 44 | 228 |
| sitemap (context) | 200 | 373 | 218 | 123 | 914 |
Three features survive excerpting. First, the emerging preference signals are rare in absolute terms: the most common, Content-Signal, is on 3.0% of parsed domains, and the other three are on single-digit counts. Against a sitemap line on 914 of 1,381 domains and a crawl-delay on 228, the AI-usage preference vocabulary is a thin layer, not an established convention. Second, three of the four emerging signals are confined to a single sector each: all 7 x-rsl and all 3 license fields are on news domains, and all 3 content-usage fields are on e-commerce domains. Only Content-Signal appears in every sector. The signal layer is not uniform across sectors; it is a few small, sector-specific pockets plus one signal with broader but still sparse reach. Third, more than one vocabulary is in play at once — Content-Signal, x-rsl, content-usage, and license are different fields expressing loosely related intents — which is itself a description of an unsettled space, stated here as a count, not as a judgment of which vocabulary is preferable.
A related malformation, not a preference signal, appears alongside these: 10 news domains carry a comma inside a user-agent value (for example User-agent: GPTBot, CCBot), a construct a conformant parser treats as a single unmatched token. It is counted here only because the same line scan surfaces it; it is analyzed in the comma-in-user-agent note and is not an AI-usage preference.
4. Authorship of the Content-Signal layer
The Content-Signal count of 42 is the largest of the emerging signals, which makes its composition the most informative single result in this census. Of the 42 carriers, 33 (78.6%) are the CDN managed default and 9 are hand-authored (Figure 1). The 33 are byte-identical: every one sets search=yes, ai-train=no, use=reference and carries the managed-comment signature. The 9 hand-authored carriers, by contrast, vary among themselves — their normalized values include search=yes, ai-input=yes, ai-train=no (3 domains), ai-train=no, search=yes, ai-input=yes (3), search=yes, ai-train=no, ai-input=no (1), search=yes, ai-train=no (1), and one setting ai-train=yes — and they stay within the policy’s three announced keys, without the use key the managed default adds [1][2]. The hand-authored files are spread as universities 1, news 3, e-commerce 5, government 0; the managed default as universities 9, news 13, e-commerce 3, government 8.
The distinction the figure draws is authorship, and it is deliberately not a distinction by CDN presence. 41 of the 42 carriers sit behind that CDN by Server header (Server: cloudflare), and that includes 8 of the 9 hand-authored files (the ninth reports apache). So the 9 are not “the domains not on the CDN”; they are domains that mostly are on the CDN but whose Content-Signal line lacks the managed-comment signature and carries an operator-chosen value. Splitting the 42 by Server header would assign 41 to one bucket and tell the reader nothing about who wrote the lines; splitting by the managed signature plus the exact default value separates the 33 injected defaults from the 9 authored lines. That is why the classification uses the signature, and why the headline figure is 33-versus-9, not 41-versus-1.
One more measured fact sits under the authorship split: the Content-Signal line is layered on blocking, not offered in place of it. 40 of the 42 carriers also root-block at least one AI crawler token in the same robots.txt file — the per-token exclusion that BA-D-2026-01 measures — so on all but two carriers the purpose-signal line and a crawler-specific Disallow coexist. Whatever a Content-Signal line expresses, in this corpus it almost always accompanies an explicit block rather than replacing one.
5. CDN fingerprint is a lower bound
The authorship result rests on identifying a CDN’s managed default, so the census also records what the corpus’s Server headers say about CDN presence — and states clearly why that record undercounts. Across the 1,381 parsed domains the Server header reads cloudflare on 288, akamai on 8, none (no Server header) on 299, and something else on 786. That last bucket is the problem for any penetration estimate: it contains CDNs that do not consistently self-identify in the Server header. Observed values inside it include cloudfront and varnish, and a fronting CDN can be entirely absent from the header, so a domain served through such infrastructure lands in other or none rather than in a CDN bucket. The share of CDN-fronted domains counted from the Server header is therefore a lower bound, and the true share is higher by an unknown margin.
Two consequences follow, and both are floors rather than point estimates. The managed-default count of 33 is a floor: it counts only carriers whose managed-comment signature and default value are both present in a readable file, and any managed default sitting behind a challenge page (Section 2.1) or under a non-self-identifying CDN is not counted. And the selection noted in Section 2.1 pushes the same direction — the unparseable_blocked domains excluded from the parsed denominator are disproportionately the heavy-CDN, bot-challenge domains, so their absence removes exactly the population most likely to carry a managed default. Every CDN-related number in this census should be read as “at least this many,” not “this many.” The carrier-side fingerprint is more reliable in the narrow case that matters for authorship: 41 of the 42 Content-Signal carriers report Server: cloudflare directly, so their CDN attribution does not depend on the other bucket — but the managed-versus-hand split among them still rests on the comment signature, not the header (Section 4).
6. Reproducibility appendix
The census is reproducible offline from the parent corpus’s stored bytes; it performs no network I/O. Inputs: the stored robots.txt responses and response headers of BA-D-2026-01 (the robots_results.jsonl behind the crawler-access dataset), keyed by (sector, domain); the parsed subset is the 1,381 rows with that census’s parsed category. Scan: for each parsed file, a raw-text line scan records the presence of a Content-Signal, x-rsl, content-usage, and license field; for Content-Signal it records the literal value and a normalized value (whitespace and key-order normalized); it records the CDN’s managed-robots.txt comment signature as a boolean; it maps the stored Server response header to one of cloudflare / akamai / other / none; and it records the crawl-delay and sitemap presence and the set of AI crawler tokens the file root-blocks (reused from the parent census’s evaluation). Authorship: a carrier is the CDN managed default when the managed signature is present and the normalized value equals search=yes,ai-train=no,use=reference; otherwise hand-authored. Every published figure is a count over the 1,381 parsed rows, and the per-domain values are in the dataset; the collection’s AGGREGATES.md, VERIFY-recompute.md, and LIMITATIONS.md record the aggregation, an independent recomputation, and the caveats. The dataset CSV was generated from the source JSONL by a deterministic serializer and re-verified cell-for-cell against both the JSONL and the published aggregates before release.
Limitations
Line scan, not a parser. Every signal is a raw-text line match, not an RFC 9309 [3] parse. The census does not resolve directive grouping, precedence, or wildcard expansion, and does not judge whether a Content-Signal, x-rsl, content-usage, or license value is semantically valid. A count is “this line is present,” not “this preference is in force.”
Single-day snapshot. The corpus is a single 2026-07-09 fetch reused unchanged; robots.txt files and their embedded signals change without notice. Every figure is a time slice and should be checked against a current fetch before it is relied upon.
Deployment, not compliance. The data shows that an operator or a CDN deployed a signal, not that any AI crawler honors it. Content-Signal, x-rsl, content-usage, and license are emerging, non-mandatory fields with no broad enforcement guarantee; nothing here measures whether any crawler acts on them, and no AI-visibility or citation effect is implied.
Server-header fingerprint is incomplete. CDN presence is read from the Server response header only, mapped to four buckets. CDNs that do not self-identify in that header (the other/none buckets include cloudfront, varnish, and fully fronted origins) are undercounted, so every CDN-penetration and managed-default figure is a lower bound (Section 5).
Parsed denominator is selective. Figures are computed over the 1,381 domains whose robots.txt parsed cleanly; the 383 unparseable_blocked, 154 fetch_error, and 82 no_robots_file domains of the parent corpus are excluded. The unparseable_blocked domains — many behind a bot-challenge page — are plausibly heavy CDN users, so their exclusion removes the population most likely to carry a managed default and biases the CDN-related counts downward. A “0” in the matrix means “a readable file without this line,” not “no file.”
Managed-signature dependence. The authorship split depends on matching a known managed-robots.txt comment signature. If the CDN changes its managed comment text, the older signature would miss newer managed files, which would move carriers from the managed-default bucket toward hand-authored and understate the managed share — again in the direction of a floor, not a ceiling.
References
- 1.Cloudflare (W. Allen). Giving users choice with Cloudflare's new Content Signals Policy (2025). https://blog.cloudflare.com/content-signals-policy/ Accessed 2026-07-10. [archived]
- 2.Cloudflare (Cloudflare Docs). Managed robots.txt — content signals and the content-use extension (2026). https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/ Accessed 2026-07-10. [archived]
- 3.M. Koster, G. Illyes, H. Zeller, and L. Sassman, IETF. RFC 9309: Robots Exclusion Protocol (2022). https://www.rfc-editor.org/rfc/rfc9309.html Accessed 2026-07-08. [archived]
- 4.IETF AI Preferences (aipref) Working Group. AI Preferences (aipref) — About / Charter (2025). https://datatracker.ietf.org/wg/aipref/about/ Accessed 2026-07-10. [archived]
How to cite
PDF of recordBarkhausen AI (2026). Emerging AI-usage preference signals in robots.txt: a signal-layer census of 1,381 domains. https://barkhausen.ai/research/ai-usage-signals-census-2026/
BibTeX
@techreport{BA-D-2026-08,
author = {{Barkhausen AI}},
title = {Emerging AI-usage preference signals in robots.txt: a signal-layer census of 1,381 domains},
institution = {Barkhausen AI},
year = {2026},
url = {https://barkhausen.ai/research/ai-usage-signals-census-2026/}
}Published under the Creative Commons Attribution 4.0 International (CC-BY-4.0).
