A 2026 preprint proposes measuring generative engine optimization as two stages rather than one count: citation selection, whether a platform triggers search and chooses a source at all, and citation absorption, how much a chosen source shapes the generated answer's language, evidence, and structure. Analyzing a self-collected public dataset of 602 controlled prompts across ChatGPT, Google AI Overview/Gemini, and Perplexity, it reports that breadth and depth diverge: Perplexity and Google cite more sources per prompt, while ChatGPT cites fewer but shows higher average per-source influence among fetched pages. The paper is an independent, non-peer-reviewed preprint built on descriptive statistics from its own data, and it says so, restricting its own claims to counts and contrasts. This note reports what the framework separates, what evidence stands behind it, and how its selection/absorption distinction relates to — without redefining — the mention/citation distinction and the metric estimands this site already maintains.
A single tally of how often a source is cited answers two different questions at once, and a 2026 preprint argues they should be measured apart [1]. It names the first outcome citation selection — whether a generative engine triggers a search and chooses a page into its citation set at all — and the second citation absorption — how much a chosen page contributes language, evidence, structure, or factual support to the answer that is generated. The distinction matters because the same observed citation can be a weak navigational reference or the source of several answer paragraphs. This note reports what the paper separates, what kind of evidence supports it, and how the distinction sits beside the vocabulary this site already maintains.
What the framework separates
The paper measures selection as a prompt-level citation count and absorption as a page-level influence score, a constructed proxy that rewards a cited page for being referenced repeatedly, appearing early, spanning more of the answer’s paragraphs, and overlapping the answer text in TF-IDF and n-gram terms [1]. Across 602 controlled prompts on three platforms, the two measures diverge. Mean citations per prompt run 6.88 for ChatGPT, 12.06 for Google AI Overview/Gemini, and 16.35 for Perplexity, so Perplexity and Google cite the most broadly; but mean influence among successfully fetched pages runs 0.2713 for ChatGPT against 0.0584 for Google and 0.0646 for Perplexity, so the sparsest citer uses each cited source the most heavily [1]. The paper reads this as evidence that breadth and depth are separate outcomes a single visibility number would collapse. A related negative result is that pages formatted as question-and-answer show marginally lower mean influence than non-Q&A pages (0.0947 versus 0.1005), which the paper treats as evidence that a surface format is not itself an absorption signal [1].
What kind of evidence this is
The framework rests on an independent, non-peer-reviewed preprint whose empirical base is a public dataset the authors assembled themselves — the geo-citation-lab repository — rather than an external or refereed source [1]. The paper is explicit about the limits this imposes. It reports descriptive statistics only, with no p-values, confidence intervals, or regression coefficients, and it states that its influence score is an observational proxy, not a measure of a model’s internal attention. It further warns that the score’s own components must not be reused as predictors of the score, which would be circular. A claim-level self-audit sorts its statements into a four-level identification map: direct counts and descriptive contrasts are treated as findings; the interpretive “evidence-container” reading is offered as an explanation consistent with the data; and causal optimization prescriptions are explicitly reserved for future intervention experiments. Read on those terms, the contribution is a measurement vocabulary and a set of reproducible descriptive patterns, not a demonstrated account of what causes a page to be absorbed.
Where it meets this site’s vocabulary
The selection/absorption cut is a different axis from the distinctions this site fixes, but it echoes their motivation. The terminology convention (BA-C-7) separates a mention — an entity named in an answer’s text — from a citation — a source attribution the answer provides — and notes that the two move independently, so a reported rate is uninterpretable until it says which it counts. The paper makes an analogous separation one level down, inside the citation itself: being selected and being absorbed also move independently, as its platform contrasts show. The visibility-metrics convention (BA-C-2), in turn, requires every estimate to name its estimand precisely, and it already carries citation-support measures alongside its mention-based primary metric. Whether that family should adopt an explicit selection-versus-absorption cross-reference — an answer-level dependent variable stated beside the mention-based one — is a question the paper’s evidence raises without settling, and one worth weighing rather than importing wholesale. Its own caution applies to any such adoption: an absorption measure is only as trustworthy as the proxy that stands in for it, and a constructed proxy must not quietly become its own explanation.
Limitations
The paper is a static snapshot. It provides no unified record-level timestamps and flags this itself, so its platform figures describe the three engines only as they behaved when the data was collected; engines change without notice, and the specific values should be assumed perishable. Its 602 prompts are designed, not randomly sampled from real traffic, and its identified sources skew heavily to US and English pages, so external validity is bounded to that prompt distribution. The absorption result depends entirely on one constructed proxy, and absorption is computed only over pages that were successfully fetched, so it is conditional on fetchability. This note reports the paper’s framework and evidence; it does not adopt selection and absorption as site terminology, and the resonance drawn with the conventions above is a proposal to weigh, not a redefinition of either.
References
- 1.Zhang Kai, He Xinyue, and Yao Jingang (independent researchers); arXiv:2604.25707v2 [cs.IR]. From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms (2026). https://arxiv.org/abs/2604.25707 Accessed 2026-07-10. [archived]
How to cite
PDF of recordBarkhausen AI (2026). Citation selection and citation absorption: what a 2026 measurement paper separates. https://barkhausen.ai/notes/citation-selection-and-absorption/
BibTeX
@techreport{citation-selection-and-absorption,
author = {{Barkhausen AI}},
title = {Citation selection and citation absorption: what a 2026 measurement paper separates},
institution = {Barkhausen AI},
year = {2026},
url = {https://barkhausen.ai/notes/citation-selection-and-absorption/}
}Published under the Creative Commons Attribution 4.0 International (CC-BY-4.0).
