← Back to Insights

June 3, 2026 · 11 min read

The evidence triad: why citations, statistics, and quotes beat keywords in AI search

The largest controlled GEO study to date tested nine optimization methods across 10,000 queries — and three evidence tactics won decisively while keyword stuffing failed. Here is what changed, why it transfers, and how to rewrite a page for AI citation.

GEO Research Content Strategy AIEO
Junwei Yu Junwei Yu CEO

If you want AI search engines to cite your content, add named statistics, embed quotations from authoritative sources, and cite third-party references — in that order of priority. In the largest controlled study of generative engine optimization to date, the Princeton GEO paper (Aggarwal et al., arXiv 2311.09735, KDD 2024) tested nine methods on a benchmark of 10,000 queries and found this “evidence triad” lifted source visibility by up to ~40%. Meanwhile keyword stuffing — the dominant SEO tactic of the last twenty years — produced little or no gain, and often did worse than the baseline. That single contrast is the most important thing a content team can internalize about AI search in 2026.

The rest of this article unpacks the experiment, explains why evidence wins and keywords fail, and turns the finding into a concrete rewrite checklist you can apply this afternoon. It closes with the part most teams skip: how to measure whether the rewrite actually moved your citations.


What the Princeton study actually measured

Most “GEO advice” circulating online is anecdote dressed as principle. The Princeton work is different because it is a controlled experiment with a metric designed for the generative era.

Traditional SEO measures rank — were you in position 1, 3, or 9? But in an AI answer, “rank” barely exists. A response is a synthesized paragraph that may weave your content into the middle of a sentence, mention you twice, or paraphrase you without a visible link. So the authors defined two new metrics:

  • Position-Adjusted Word Count (PAWC) — how much of the generated answer is attributable to your source, weighted by where it appears (earlier and more prominent mentions count for more).
  • Subjective Impression — an LLM-judged sense of how influential and authoritative your source feels within the answer.

They then took real web pages, applied nine distinct rewriting strategies to each, regenerated answers across the benchmark, and measured the lift. Crucially, the queries span domains — debate, science, how-to, history, commercial — so the result is not a single-niche fluke.

The headline is not “GEO works.” It is that GEO methods are not interchangeable: the gap between the best and worst tactic is the difference between a +30–40% visibility gain and negative movement. Choosing the wrong playbook is not merely ineffective — it can actively suppress you.


The nine methods, ranked by effect

Here is the practical core of the paper — the methods tested and roughly how each moved Position-Adjusted Word Count. The three winners share a single quality: they add verifiable evidence the model can lift and attribute.

MethodWhat it doesEffect on visibility (PAWC)
Cite SourcesAdd citations to credible third-party references+30–40%
Quotation AdditionEmbed direct quotes from experts / authoritative sources~+41%
Statistics AdditionInsert relevant, named quantitative data~+37%
Fluency OptimizationImprove readability and flowModest positive
Easy-to-UnderstandSimplify languageModest / mixed
Authoritative toneRewrite to sound more confidentModest / mixed
Technical TermsAdd domain jargonMixed
Unique WordsInsert distinctive vocabularyMarginal
Keyword StuffingRepeat target search keywordsLittle/none — often worse than baseline

Three observations sit on top of this table.

First, the evidence triad — Cite Sources, Quotation Addition, Statistics Addition — is in a class of its own. These are not stylistic tweaks; they change the epistemic substance of the page. A generative model assembling an answer is, in effect, looking for sentences it can safely stand behind. A sentence carrying a named statistic or an attributed quote is safer to repeat than an unsupported assertion, because the evidence travels with it.

Second, fluency and tone help only at the margin. Making a page read nicely is table stakes — it does not differentiate you, because everyone else’s page reads nicely too.

Third, keyword stuffing does not transfer from SEO to GEO. This is the finding that should reorganize budgets. The single most-practiced SEO tactic of the last two decades — densely repeating the query terms — is inert at best and counterproductive at worst in AI answers. Generative engines are not counting term frequency; they are evaluating whether a passage is quotable and defensible.


Why evidence wins (and keywords don’t)

The mechanism is worth understanding, because it tells you what to do in cases the paper didn’t test.

A classic search index ranks documents largely by lexical and link signals — does the page contain the query terms, do authoritative pages link to it. Repeating keywords was rational under that regime. A generative engine works differently: it reads candidate passages and decides which to incorporate into a fluent, attributable answer. The selection pressure is no longer “does this match the query string” but “can I quote this and have it hold up.”

Under that pressure:

  • A named statistic is a unit of borrowed credibility. “Adoption rose 37% year over year (source X)” is something the model can drop into an answer with a citation. “Adoption rose significantly” is not.
  • A quotation is pre-attributed authority. A direct quote from a recognized expert lets the model cite the expert through you — your page becomes the conduit, and it gets credited.
  • A third-party citation signals the page is part of a verifiable web, not a closed assertion. It lowers the model’s risk of repeating something unsupported.
  • A repeated keyword carries no evidence at all. It adds nothing the model can attribute, so it is invisible to the selection mechanism — and density that trips spam heuristics can push you below baseline.

This reframes GEO writing from “optimize for a ranking algorithm” to “supply the model with sentences it would be comfortable quoting under its own name.


The challenger advantage: who gains the most

The single most actionable nuance in the paper is that the effect is not uniform — it depends heavily on where you start.

Lower-ranked sources gain the most. A page sitting around fifth position saw gains as large as ~+115% from Cite Sources alone. The intuition: a generative engine is hungry for usable evidence and will reach further down the candidate list to find a passage that gives it a quotable statistic or an attributed claim. Evidence lets an underdog page punch above its rank.

Conversely, already-dominant sources can lose ground — by as much as ~30% in some conditions. When you are already the default citation, adding more evidence can dilute your share as the engine now has other well-supported passages to choose from, and it spreads attribution around.

For a challenger brand in a crowded Japanese B2B SaaS category, this is the most encouraging result in the GEO literature: the evidence triad is structurally biased toward the under-cited. If you are not yet the default answer, evidence is your highest-leverage move.

The strategic read: incumbents should defend with consistency and authority breadth; challengers should attack with evidence density. They are not the same game.


A concrete rewrite checklist

Translate the findings into a page-level checklist. Run an existing article through these four moves before you touch anything else.

1. Answer first. Put the direct, extractable answer in the opening — ideally the first 50 words, before any preamble. This is both intuitive and grounded: analyses of LLM-cited content find that roughly 44% of citations come from the first 30% of a page. Front-load the conclusion; the model reads the top of the page most carefully.

2. Add named statistics. Replace every vague quantifier (“many,” “rapidly,” “significantly”) with a specific number and its source. “Faster adoption” becomes “adoption grew 37% in 2025 (per [named report]).” Each one is a unit the model can lift verbatim with attribution.

3. Embed authoritative quotes. Insert at least one direct quotation from a recognized expert, a primary study, or an industry body — with attribution. You are handing the engine pre-credentialed authority it can pass through your page.

4. Cite third-party sources. Link out to credible references for your factual claims. Counterintuitively for SEO veterans, linking out raises your odds of being cited, because it places your page inside a verifiable evidence web rather than presenting it as an unsupported island.

What is conspicuously absent from this list: stuffing the page with your target keywords. The data says that work is, at best, wasted.


Where structure meets evidence: GEO-SFE

The Princeton study isolates the evidence dimension — what a page says and what it cites. Our own research isolates the complementary dimension: how a page is structured. In “Structural Feature Engineering for Generative Engine Optimization” (arXiv 2603.29979), we showed that structural rewrites alone — sectioning, fact placement, list use, visual emphasis — lifted citation rate by +17.3% and subjective answer quality by +18.5% across six engines, holding the meaning constant.

These two findings stack rather than compete. Evidence gives the model something worth quoting; structure makes that evidence easy to find and lift. A named statistic buried in a wall of text mid-page underperforms the same statistic placed in an answer-first opening, in a short paragraph, with the number bolded. GEO-SFE also accounts for the fact that engine architectures differ — Gemini (an STS-style synthesizer), Perplexity (an IR-first retriever), and ChatGPT / Claude (in-context selectors) weight structure differently — so the optimal placement of your evidence is engine-aware, not one-size-fits-all.

This is exactly the layer VeReach GEO operationalizes. The evidence checklist above tells you what to write; the GEO-SFE structural layer tells you where to put it so each of ChatGPT, Gemini, Claude, and Perplexity is most likely to lift it.


How VeReach GEO measures the lift

A rewrite is a hypothesis. The only way to know it worked is to measure citation behavior before and after — which is harder than it sounds, because AI citations swing month to month and differ across engines.

VeReach GEO closes that loop:

  • Visibility Score (VS = 0.4·Coverage + 0.3·Position + 0.3·Influence) turns scattered mentions into one comparable number, so a rewrite’s effect is legible rather than anecdotal.
  • Four-engine tracking (ChatGPT / Gemini / Claude / Perplexity) shows whether your evidence transfers across architectures or only lands on one — the same cross-engine generalization our paper cares about.
  • A Japanese four-tier source dictionary weights citations by domain authority — ITmedia / @IT / CNET at 5.0, BOXIL / ITreview / IT Trend at 3.0, MarkeZine / PR TIMES at 1.5, note / Zenn / Qiita at 1.0 — so “cite third-party sources” becomes a quantified tactic, not a vibe. Quoting a tier-1 source is measurably worth more than quoting a tier-4 one.
  • Brand-name normalization (katakana / kanji / romaji / ASCII) ensures a mention counts whether the engine writes your name in any of its Japanese forms — essential for honest measurement in the JP market.
  • Per-Focus time-series (each Focus tracks one archetype) tracks the rewrite’s effect on the specific business question you care about, over time, rather than as a one-off snapshot.

The workflow is the loop the evidence triad demands: rewrite a page with named statistics, embedded quotes, and third-party citations; structure it with GEO-SFE; then watch the Visibility Score across all four engines to confirm the lift is real and durable.


A necessary caveat

Intellectual honesty requires one qualification. The Princeton experiments were run in the GPT-3.5 era, and effects are domain-dependent. The exact magnitudes — “+41% from quotes,” “+115% for a fifth-ranked site” — should be read as directional evidence from that setting, not as guarantees on today’s frontier models. Independent industry replications point the same way (one 2026 study reports roughly +22% from statistics and +37% from quotes), which is why we trust the mechanism: generative engines select passages they can quote and attribute, and evidence makes a passage quotable. That logic has, if anything, strengthened as engines lean harder on grounding and citation.

So: adopt the evidence triad as principle, abandon keyword stuffing, and verify the magnitude in your own domain with continuous measurement. Principle plus measurement is the whole discipline.


Want to see which of your pages the AI engines actually cite — and how much an evidence-and-structure rewrite moves your Visibility Score across ChatGPT, Gemini, Claude, and Perplexity? See how VeReach GEO measures it, or talk to the team.


A note on sources: the nine-method results, the evidence-triad effects, and the challenger-advantage figures are drawn from Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024); the structural results from arXiv 2603.29979; the “44% from the first 30%” figure and the 2026 replication from public industry analyses. Figures are as of June 2026 and reflect the model generations available at the time of each study — verify against current engines before any high-stakes decision.