A 2026 paper measured how often brands at three levels of market stature were mentioned in AI answers on their first tracking run, using one vendor's own production telemetry: 102 brands, 102,025 prompt responses, and 149,912 source citations across five engines, gathered March–May 2026. Restricted to category questions that did not name the brand, first-run visibility formed a near-linear ladder — global household names 72.9% (95% CI [60.1, 84.2], n = 11), established mid-market brands 43.6% ([36.4, 50.9], n = 36), and niche or small brands 11.4% ([4.2, 20.3], n = 55) — steps of 29.3 and 32.2 percentage points, with a Kruskal–Wallis omnibus test and all three Bonferroni-corrected pairwise comparisons rejecting equality. This note reports what the study measured and reads its bottom rung against the base-rate account: an 11% first-run figure places niche brands in the near-zero regime where a single reading is least reliable and boundary-valid intervals are mandatory. It states what a single-source, single-author, single-first-run, observational design does not support.
A niche brand and a global household name do not enter an AI answer from the same starting point, and a 2026 paper put a number on the distance [1]. Working from a single vendor’s production telemetry, it measured how often brands at three levels of market stature were mentioned on their first tracking run when a category question did not name them, and found a near-linear ladder: about 73%, 44%, and 11%, roughly 30 percentage points per step [1]. This note reports what the study measured and reads its bottom rung against this publication’s account of low base rates, and states what a single-source, single-author, single-first-run design does not support.
What was measured and how
The empirical section draws on one company’s production database: 102 brands, 3,508 completed tracking runs, 102,025 prompt responses, and 149,912 source citations across five engines, gathered between March and May 2026 (§6) [1]. The headline metric is per-brand prompt-level visibility on each brand’s first completed tracking run, restricted to unbranded category prompts, where the brand name does not appear in the prompt; branded prompts, which recognize the named brand by construction at 94% or higher on every engine, are excluded from it (§6.1) [1]. Each brand was assigned to a tier deterministically from a hand-coded rubric of public stature signals — Wikipedia coverage, press, and funding or public status (§6.1) [1]. Uncertainty uses a nonparametric bootstrap of 10,000 resamples at the brand-cohort level, and the tier comparison uses a Kruskal–Wallis test with pairwise Bonferroni-corrected Mann–Whitney tests, chosen because the per-brand distributions are bounded and right-skewed (§4.5, §6.1) [1]. The author describes the reading as cross-sectional and observational and reports it as a quantification, not a causal claim (§6.1) [1].
The three-tier ladder
First-run unbranded visibility fell in clean steps. Global household names (Tier 1, n = 11) were mentioned in 72.9% of unbranded category answers (95% CI [60.1, 84.2]); mid-market and regional brands (Tier 2, n = 36) in 43.6% ([36.4, 50.9]); and niche or small brands (Tier 3, n = 55) in 11.4% ([4.2, 20.3]) — steps of 29.3 and 32.2 percentage points (§6.1, table) [1]. A Kruskal–Wallis test rejected equality of the tier distributions (H = 38.32, df = 2, p = 4.78 × 10⁻⁹), and all three Bonferroni-corrected pairwise comparisons at α = 0.05/3 rejected as well: Tier 1 vs Tier 2 (U = 343, p_Bonf = 8.1 × 10⁻⁴, d = 1.57), Tier 1 vs Tier 3 (U = 575, p_Bonf = 8.3 × 10⁻⁶, d = 2.34), and Tier 2 vs Tier 3 (U = 1629, p_Bonf = 6.4 × 10⁻⁷, d = 1.31), all effect sizes large by convention (§6.1) [1]. Because the Tier 1 cell is small, a leave-one-out sensitivity moved its mean only within [70.5, 76.7]%, so no single brand carries the headline (§6.1, §8) [1].
An 11% floor is a low-base-rate measurement
The bottom rung is the point of contact with the base-rate account. A cohort mean of 11.4% for niche brands, with an interval reaching down to 4.2% (§6.1) [1], places them in the near-zero regime that the convention on visibility metrics (BA-C-2) singles out: where a challenger or niche entity’s true visibility sits close to the floor, a symmetric normal interval can extend below zero — an impossible negative visibility — so a report must use a boundary-valid construction such as the Wilson score or Clopper–Pearson interval. The study’s use of a bootstrap rather than a normal approximation is consistent with that requirement, and its asymmetric Tier 3 interval is the boundary showing through. The note on base rates and regression to the mean (BA-MN-4) adds the reliability point: for an entity whose true visibility is near zero, a single reading is the least informative kind, and a portfolio scanned once will surface some chance appearances regardless of any effect. The paper’s 11.4% is a cohort mean over 55 brands, far steadier than any one brand’s number — but that is the lesson turned outward: a niche brand measured on one tracking run is a single observation in this low-base-rate region, where BA-C-2’s arithmetic says a lone reading cannot support a visibility claim. The ladder is field-scale evidence that most challenger and niche brands begin their measurement life there.
Limitations
The design bounds the reading in ways the paper itself names. The evidence is single-source and single-author: one company’s first-party production telemetry, analyzed by a member of that company, with no independent replication alongside it. The cohort is convenience-sampled and skews toward SaaS, retail-execution, fintech, and Indian direct-to-consumer brands, so it is not claimed to be category-representative (§8) [1]. The headline is a single first-run cross-section, not a repeated-measures design, so it speaks to where brands start rather than how they move, and it inherits the run-to-run instability documented elsewhere for these engines. The Tier 1 cell is small at n = 11, and the tier assignment is a hand-coded proxy for web prominence whose inter-rater reliability is not yet quantified (§6.1, §8) [1]. Each engine was measured in one search-enabled variant held fixed, so whether the ladder reproduces on other configurations is untested (§8) [1]. The figures are descriptive statistics of an observational dataset, and the paper is explicit that they support no causal claim (§7.2, §8) [1]; and they describe five engines over one March–May 2026 window and should be assumed perishable. Ranqo appears here only as the author’s affiliation and the origin of the data; nothing in this note evaluates or endorses it as a tool.
References
- 1.Kumar. Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines (2026). https://arxiv.org/abs/2606.20065 Accessed 2026-07-10. [archived]
How to cite
PDF of recordBarkhausen AI (2026). The brand-stature visibility ladder: what one 2026 first-party study measured. https://barkhausen.ai/notes/geo-at-scale-brand-tiers/
BibTeX
@techreport{geo-at-scale-brand-tiers,
author = {{Barkhausen AI}},
title = {The brand-stature visibility ladder: what one 2026 first-party study measured},
institution = {Barkhausen AI},
year = {2026},
url = {https://barkhausen.ai/notes/geo-at-scale-brand-tiers/}
}Published under the Creative Commons Attribution 4.0 International (CC-BY-4.0).
