The Research Behind AI Citation: What the GEO Papers Actually Found
Adding concrete numerical data to a page raised AI source visibility by up to 40% in controlled testing. That finding — from Aggarwal et al. at ACM KDD 2024 — is the most actionable single result in the generative engine optimization literature. But it is one of four GEO papers that, taken together, establish something more specific than "add more data." They establish exactly which content properties matter, why they matter, and why citation authority accumulates rather than resets.
The full citations and capsules are in the AI Citation & GEO research index. This post covers what each paper found and what it means for a marketing team trying to get cited.
The foundational GEO study: what content modifications actually work
Aggarwal et al. (2024, ACM KDD) ran the first systematic study of which content modifications increase a source's likelihood of appearing in AI search answers. They tested nine content strategies across five generative engines — Bing, You, Perplexity, NeevaAI, and Llama — using a controlled benchmark.
The result by strategy:
- Statistics Addition (adding concrete numerical data): up to 40% visibility improvement
- Adding citations and quotations: measurable improvement
- Keyword stuffing: no measurable benefit
The 40% figure is an upper bound across their test matrix, not a guaranteed outcome for every page. But the directional finding is unambiguous: specific, numerical claims are what generative engines surface. Content that reads well but makes no concrete, attributable claims gets no corresponding lift.
For marketing content, this changes what specificity means. It is not a writing virtue — it is a citation signal. "Our platform reduces editing cycles" has no GEO value. "Our platform reduces editing cycles by an average of 3.2 per piece" does.
Citation selection vs. citation absorption: two separate problems
Zhang et al. (2026, arXiv 2604.25707) identified a distinction the earlier research missed: getting retrieved and actually shaping the generated answer are not the same problem.
Citation selection is whether an AI engine pulls your URL into its context at all. Citation absorption is whether your content, once retrieved, actually appears in the generated answer. A page can pass citation selection and contribute nothing to the output — the model retrieved it and ignored it.
Zhang et al. identified four traits of high-absorption pages:
- Greater length
- Structured formatting
- Semantic alignment with the query
- Extractable numerical evidence
The practical implication: if your pages are being retrieved but your brand is not appearing in AI-generated answers, the problem is absorption, not discovery. The interventions are different. Discovery is addressed through topical coverage and authority signals. Absorption is addressed through structure and extractable specificity — the same content properties Aggarwal et al. identified.
Structural optimization improves citation independently of content
Yu et al. (2026, arXiv 2603.29979) tested whether restructuring content — without changing its underlying information — improves AI citation.
They tested optimization at three levels:
- Document architecture: overall organization and hierarchy
- Information organization: how information is grouped and sequenced within sections
- Visual emphasis: headers, lists, and formatting that signal structure
Each level independently produced statistically significant improvement in both citation rate and citation quality. The finding is consequential: a page containing accurate, relevant information in dense prose will be cited less often than the same information organized with clear hierarchy, labeled sections, and formatted lists — regardless of the content itself.
This is the research basis for the practical guidance that brand context should be organized as directive lists and labeled sections, not as paragraphs. The organizational signal matters to AI extraction independent of what the content says.
Citation authority compounds: the concentration finding
Yang (2025, arXiv 2507.05301) analyzed which sources AI search systems actually cite, across query types and over time. The finding: AI citation distribution is narrow and stable. A small set of sources receives the majority of citations, and that preference pattern is consistent across queries.
The implication for marketing teams: new or low-authority sources face a structural disadvantage regardless of content quality alone. A single well-optimized page does not overcome the concentration dynamic. What the research establishes is that citation authority in AI search accumulates through consistent, structured content presence over time — not from any single piece.
This reframes GEO as compounding investment rather than campaign optimization. Each piece of well-structured, fact-dense content that AI engines can cite contributes to an accumulating citation authority position. Teams that start earlier compound longer.
What this research does not settle
The GEO literature has meaningful gaps. Most studies use simplified content scenarios and controlled benchmarks; field studies with real marketing content and real search queries are limited. The 40% visibility improvement figure comes from a specific benchmark matrix, not a representative sample of marketing content performance.
What the research does establish with consistency: structure and specificity are the primary content-level citation signals, keyword density is not, and consistent citation presence compounds in a concentrated distribution. Those three findings hold across the literature.
For SEO vs. GEO differences, structuring content for AI citation, what GEO is, and why AI cites some brands and not others — those posts cover the applied implementation side.
Frequently Asked Questions
Does the 40% visibility improvement apply to all content?
No. The Aggarwal et al. finding came from a controlled benchmark across five generative engines with specific content scenarios. The improvement is directional evidence that Statistics Addition is an effective GEO strategy; it is not a guarantee of a specific outcome for any marketing page. The finding should be read as: "concrete numerical data outperforms the alternatives" — not as a predicted lift percentage for a specific implementation.
What is the difference between GEO and traditional SEO at the content level?
SEO optimizes for a human choosing from a ranked list of results; the content needs to deliver on the heading promise and keep the reader on page. GEO optimizes for a model extracting a passage that can stand alone in a generated answer; the content needs to be extractable, self-contained, and attributable without surrounding context. The structural changes that GEO requires — answer-first paragraphs, FAQ schema, concrete numerical claims — are incremental to SEO, not in conflict with it.
If citation is concentrated in established sources, how does a smaller brand compete?
The Yang concentration finding is about overall distribution, not absolute barriers. The GEO research also establishes that content-level signals — specific numerical claims, structured formatting, sourced statements — improve citation independently of site authority. A smaller brand publishing a single well-structured, fact-dense page on a specific topic can outperform a larger brand's thin treatment of the same topic in AI citation for that query. The concentration dynamic makes consistent, quality-first content publishing more important, not less.
How long before GEO changes appear in AI citations?
The research does not establish a standard lag time. Generative engine citation behavior reflects what the model has indexed and can access at retrieval time. Changes to publicly accessible content can affect AI citations within days for frequently queried topics. The consistent finding is that structural changes produce measurable improvements; the timeline depends on how often the relevant topic is queried and how frequently the generative engine refreshes its index.