Most GEO Advice Does Not Replicate
Every agency is selling a Generative Engine Optimisation checklist. A survey of 45 studies found that almost none of it survives a controlled test.
In this article
The gap between what is sold and what has been testedRetrieval versus citation, the distinction that breaks most auditsWhat the evidence actually supports, with the conditions attachedWhat the evidence does not support, naming the popular tacticsWhy the famous statistics finding is much narrower than the headlineDrift, and why any GEO finding has a short shelf lifeWhat ‘we optimised for AI search’ should mean in a scope of workHow to run your own test instead of trusting a checklistWhat to measure when citation-to-click is unprovenHow to have this conversation with a client, honestlyThe gap between what is sold and what has been tested
Open any Indian agency’s GEO deck from this year and you will find roughly the same twelve bullets. Add statistics. Add quotations. Add schema. Write in a question-and-answer format. Publish an llms.txt. Get cited on Reddit. Keep paragraphs short so the model can extract them. The deck presents these as settled practice, the way canonical tags or alt text are settled practice, and it usually attaches a percentage to at least one of them. The bullets rarely carry a citation.
Very little of that has been tested. In July 2026 Olivier Martinez published a critical scoping review of 45 GEO studies covering November 2023 to July 2026, drawn from arXiv, the ACM Digital Library, the ACL Anthology and adjacent databases, and the picture it assembles is far bleaker than any deck admits. The survey’s own summary is that already-retrieved content can causally influence an answer, and that no technique showed a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream clicks and conversions.
Read that twice. It is not saying the tactics are wrong. It is saying that for most of them nobody has yet run the experiment that would settle the question either way, and that in the cases where somebody did run it properly, against a real retrieval pipeline rather than a fixed context, the result came back null. Absence of evidence, mostly.
There is an honest reason the decks look the way they do. The field is about two and a half years old, it moves faster than publication cycles, and a client who asks ‘what should we do about ChatGPT’ wants a list, not a literature review. The dishonest part is the confidence. A tactic that has been tested once, in a fixed context, on one engine, with a single task type, is being sold as a discipline with established best practices and a service-line price attached, and the gap between what was measured and what is promised is where the budget quietly goes. That gap is the subject of this post.
Retrieval versus citation, the distinction that breaks most audits
Visibility is not one thing. That is the single most useful idea in the survey, and it is the one most decks skip. Martinez decomposes it into a vector of seven separate quantities. Three matter here. Retrieval probability comes first, meaning whether your page is pulled into the model’s working set at all. Context exposure comes next: where in that set it lands, and how many tokens it is given. Then citation probability, meaning whether the answer names you.
These are different stages. A tactic can help one and hurt another.
Almost every AI-search audit sold in India collapses all three into a single score, usually built by prompting a handful of engines with a few brand-relevant questions and counting how often the brand appears. That number is a citation count. It tells you nothing about whether the failure sits at retrieval, at ranking, or at generation, and those three failures have completely different remedies. If you are never retrieved, then no amount of rewriting the body copy will help you at all, for the simple reason that the model never reads the body copy of a document it did not pull into the working set. If you are retrieved and never cited, rewriting is exactly right. An audit that cannot tell you which situation you are in is not an audit, it is a tally. Ask which stage failed.
The survey’s pipeline model makes the sequence explicit: activation, then retrieval, then ranking, then generation, then whatever the user does next. Each stage is stochastic and each is only partly observable from outside. You can see the final answer. You cannot see the retrieved set, the reranked order, or the token budget, which means most of the causal chain is hidden from anyone without engine access, and inferences about it are inferences rather than measurements.
Keep the stages separate in your own reporting and a lot of confusion dissolves. Merge them and you will spend a quarter optimising the wrong stage. Separate the three.
What the evidence actually supports, with the conditions attached
Two levers survive. Both carry strong support in the survey’s confidence table, and both are unglamorous enough that no agency has ever built a deck slide around either of them.
The first is query-document relevance. Content that genuinely and specifically answers the question gets cited more, which is not a surprise, and the survey grades this as high confidence on the strength of counterfactual experiments. Wan and colleagues found in 2024 that topical relevance dominates over signals like scientific references or a neutral tone, which is a quiet rebuke to a lot of the tone-and-formatting advice in circulation. Say the useful thing. Say it where the question is asked, in the words the question uses, and the largest single lever in the published evidence is already working in your favour before anyone touches a formatting rule.
The second is position within the context window. Where your document sits in the set handed to the generator strongly predicts whether it gets cited. The strongest single piece of evidence here is a factorial study by Vishwakarma and colleagues running 252,000 trials across six large language models and eighteen factors, which concluded that topical relevance and context position are the primary determinants of the first citation. That is a large, well-powered design. Its conclusion points squarely at the two things a content team has the least direct control over.
Note the condition. Position in the context window sits downstream of retrieval, which means the effect only exists for documents that have already been pulled in, and a page that never enters the working set cannot benefit from it by any amount of editing. It was conditional.
Below those two sits a middle tier. These are the moderately supported levers, useful and heavily conditioned. Extractable evidence, meaning claims that can be lifted cleanly and attributed, shows moderate to strong support with a truthfulness constraint attached. Recency, explicit prices and dates show moderate support, but only for time-sensitive queries, which makes them a targeted tool rather than a general one. Document structure shows moderate and heterogeneous support, with effects that are stage-dependent, helping retrieval in some setups while doing nothing for citation. Fluency and simplification are weak to moderate and vary by domain and engine.
What the evidence does not support, naming the popular tactics
Start with the clearest result. Keyword stuffing does not work in generative engines, it is graded null or negative across multiple benchmarks, and the survey’s recommendation is simply to avoid it. The tactic does not transfer from conventional SEO. Anyone still selling it is selling 2011. Drop it entirely.
The more important negative result is C-SEO Bench, by Puerto and colleagues, because it tested the checklist approach directly and at scale. Across two tasks, six domains, roughly 1,900 queries and 16,360 documents, only three of 54 method-and-domain combinations were significantly positive in the main experiment. None was positive on question answering. That is close to what you would expect from chance alone, and question answering happens to be precisely the task that AI search consists of, which means the benchmark returned a null result on the only task type most clients are actually paying to influence. Three out of fifty-four.
The same benchmark found worse news. It is worse specifically for the agency pitch. Gains decline as adoption increases, producing what the authors describe as congested dynamics. If a tactic works because few people use it, selling it to everyone destroys it, which is an uncomfortable property for a service line built on repeatable playbooks.
Then there is SAGEO Arena, which tested optimisation against a realistic retrieval pipeline rather than a fixed context. Across 171,003 documents and 2,700 queries, body-only optimisation reduced average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by 6%. Optimising the body text alone made things worse at every stage measured. That is the clearest demonstration available that a tactic validated inside a fixed laboratory context can invert entirely once retrieval is allowed to do its job, which is the condition every real page on the open web actually operates under. The finding does not generalise upward.
Two more absences are worth stating carefully. The survey does not treat schema markup or llms.txt as tested levers at all, which is not evidence that they fail, only that the research base has not examined them and that confident claims about them rest on nothing published. And on earned media, Chen and colleagues observed in 2025 that earned coverage is overrepresented among cited sources, but that is an observational finding. Overrepresentation is not causation, and no mechanism has been shown by which third-party coverage mechanically causes a recommendation.
Why the famous statistics finding is much narrower than the headline
The number that launched the field came from the Princeton GEO paper, Aggarwal and colleagues in 2024, and it is usually quoted as ‘GEO lifts visibility by up to 40%’. The survey traces it. Provenance changes what the figure means, and the trace is worth following carefully, because almost every downstream repetition of the number has dropped the condition that makes it true.
Start with the metric. It is position-adjusted word count, which discounts text appearing later in an answer. For the Quotation Addition treatment, that metric moved from 19.3 to 27.2, or roughly 41% in relative terms. So the number is real arithmetic on a real result. It is also a relative maximum on one metric in one configuration. The framing is the problem, not the arithmetic.
Here is the condition that matters. The five documents were already provided to the generator in a fixed context. Retrieval was not part of the experiment. As the survey puts it, the result does not mean wider discoverability, it means that in this testbed a source already provided receives a larger position-weighted share of the answer. Discoverability was held constant by design, which means the study could not have detected an effect on it either way.
Translated into practice: the finding says that if a model is already reading your page, adding a well-placed statistic can win you more of the answer. It says nothing whatsoever about how to get the model to read your page in the first place, which happens to be the problem almost every client actually walks in with, and which the experimental design deliberately removed in order to isolate the effect it wanted to study. Different question entirely.
The survey grades the claim that GEO increases visibility by 40% as rejected. Not weak, not unproven. Rejected, on the grounds that it is a relative maximum on a single metric under a specific configuration, presented as a general effect. If a proposal on your desk quotes that number without the fixed-context caveat, the person who wrote it has not read the paper they are citing.
Drift, and why any GEO finding has a short shelf life
Even a correct finding decays. The survey treats platform volatility as a structural problem rather than a nuisance, and the measurements it assembles are sobering enough that they should change how any practitioner reads a result more than a few months old, including the results in this post. Shelf life is short.
Schulte and colleagues reported in 2026 a Jaccard similarity of 0.34 to 0.42 across four engines over 45 days, meaning the overlap between which sources an engine cites now and which it cited six weeks ago is roughly a third to two fifths. Worse for anyone running tests, at temperature zero, where output is supposed to be deterministic, 9 to 28% of decisions changed on repeated runs. Kirsten and colleagues found 18% page overlap for AI Overviews over two months, against 45% for organic Google results, so the AI surface churns at roughly twice the rate of the blue links beneath it. Read those numbers again.
Three consequences follow. A single-run observation is noise. A finding published nine months ago is describing a system that no longer exists in that form, so its conclusions transport only as far as the engine name, model version, date and locale it was recorded against. And a vendor dashboard showing your ‘AI visibility score’ ticking up and down week to week is mostly showing you sampling variance dressed as insight, unless it runs enough repetitions per cell to separate signal from churn, which most of them do not disclose either way.
The survey’s answer is living benchmarks. It argues for versioning prompts, snapshots and content, rather than publishing scores presented as timeless. The practitioner version of the same idea is that your measurement has to be continuous and repeated, because a quarterly point estimate of something this unstable is close to meaningless.
What ‘we optimised for AI search’ should mean in a scope of work
If the phrase is going to appear on an invoice, it needs a definition that survives contact with the evidence. Here is one that does.
Four commitments, in order. It should mean, first, that the page is technically retrievable and genuinely answers a specific question that real people ask, because relevance is the best-supported lever there is. Second, that claims on the page are extractable and true, phrased so a model can lift a sentence and attribute it without distorting it. Third, that retrieval, citation and fidelity are measured separately rather than rolled into one score. Fourth, that any tactical claim beyond those is labelled as an experiment with a stated hypothesis and a review date, rather than as a deliverable with an implied outcome.
It should not mean a recipe. The survey grades fixed recipes as generalising poorly and requiring multi-engine testing before anyone trusts them, and C-SEO Bench is the direct demonstration.
Honesty is not the only argument. There is a commercial one for writing scopes this way. A scope that promises citation counts is a scope you can be held to on a metric that drifts by a third every six weeks through no fault of yours, which is a bad trade for the agency as well as the client. A scope that promises a defined body of work, a defined measurement protocol, and honest reporting of what moved is one you can actually deliver, and it survives the quarter when the engine changes its retrieval stack and everyone’s numbers fall over at once.
Write the uncertainty into the document. It reads as competence, not hedging.
How to run your own test instead of trusting a checklist
The survey sets out a minimum protocol for researchers. It adapts cleanly to agency work. The core move is to stop asking ‘does this tactic work’ and start asking ‘does this tactic work on this engine, for this query type, in this domain, this month’.
Decide first what you are estimating. A conditional effect, meaning the change given that you were retrieved, is a different question from a total effect, which includes retrieval, and both differ from a commercial effect. Pick one before you collect data. Most in-house tests fail here, because they measure a conditional effect and report it as a total one.
Then build a factorial design. Cross the engines you care about, recording the model name, mode, date, locale, account type and whether search was enabled, since the survey notes that a commercial product name can conceal several different systems. Cross the domains and analyse them separately, because pooling domains is how real effects get averaged into nothing. Use three to five paraphrases of each query intent. Users do not phrase a question the same way twice. Include an untreated baseline, the intervention, and a placebo change that should do nothing. Randomise the order of documents in context where you control it.
On sample size, the survey cites Schulte and colleagues suggesting seven to eight repetitions per cell as a reasonable starting point, validated against a pilot estimate of variance. Do the multiplication before you commit. Four engines, five paraphrases, two conditions and eight repetitions is 320 observations for a single query intent in a single domain, which is a real piece of work and considerably more than the two prompts most audits are built on.
Record more than the citation. For each run, log whether the page appeared in the retrieved set if you can see it, whether it was cited, where in the answer it landed, whether the claim attributed to you was actually yours, and whether the attribution was accurate. Fidelity matters commercially. Being cited for something you did not say is a reputational event, not a win, and a citation count cannot tell the difference.
What to measure when citation-to-click is unproven
Clients find this part hardest. It deserves a straight answer. The link between being cited in an AI answer and getting a click, a lead or a rupee of revenue has not been established.
The survey grades citation-to-conversion claims at very low confidence, resting on one suggestive quasi-experiment plus industry assertions that lack rigour. That quasi-experiment, by Watanabe and Nakayashiki in 2026, estimated an additional traffic multiplier of 1.82 with a 95% interval of 1.31 to 2.54, which sounds encouraging until you read the next clause: a conservative temporal placebo test yielded p=0.16. In plain terms, when the authors checked whether their method would find an effect in a period where there should not be one, the check did not cleanly pass. One study, one geography, one moment in a fast-moving system. Nobody has replicated it.
So build your measurement in two layers. The lower layer is what you can observe directly and repeatedly: retrieval where visible, citation rate by query intent, position in the answer, and attribution fidelity. Track those as process metrics, with confidence intervals, never as single numbers.
The upper layer is business outcome, and it needs the same discipline you would apply to any channel whose attribution is broken. Watch referral traffic from assistant domains where your analytics can see it. Watch direct and branded search volume for the queries you are targeting, since an AI answer that names you and is not clicked can still produce a branded search later. Watch lead quality and self-reported source on your enquiry forms, since a single ‘how did you hear about us’ field will catch attribution that no tracking parameter can reach, particularly for a channel where the referring surface often strips the referrer entirely. It is crude and it works. Then use difference-in-differences against a matched set of untreated pages, with a temporal placebo, rather than reading a before-and-after chart and calling it causal.
No clean number emerges. A dashboard wants one anyway. That is the honest state of the field, and pretending otherwise is how agencies end up defending a metric they cannot explain.
How to have this conversation with a client, honestly
The fear is obvious. Saying ‘we do not know yet’ is supposed to lose the pitch. In practice it usually wins it, because the client has already heard four decks full of certainty and has noticed that none of them agreed with each other.
Lead with what is solid. Relevance and context position are well supported, the content work that follows from them is work you would want done regardless of AI search, and it does not depreciate if the engines change. That is a defensible core. Then be specific about the boundary: the popular tactics have either failed to replicate, worked only in fixed-context laboratory conditions, or reduced retrieval when tested against a real pipeline, and here are the three studies that show it.
Name the open question directly. Whether AI citations produce clicks and revenue at any reliable rate is unresolved, one quasi-experiment points in a positive direction and did not survive its own placebo test cleanly, and any agency quoting you a conversion rate for AI search has made it up. Say that plainly. It reframes the entire evaluation, because the client now has a test they can apply to every other proposal in the pile.
Then give them the thing certainty cannot give them, which is a decision rule. Run the content work that holds up regardless. Treat everything else as an experiment with a hypothesis, a sample size, a review date, and permission to be abandoned when it fails. Report the failures. A client who sees you kill your own tactic on evidence will believe you the next time you say something worked, and that credit is worth more over a two-year relationship than any single quarter of confident numbers.
The field will firm up. Better designs are being published, the survey itself sets out what those should look like, and in eighteen months some of today’s open questions will have answers. Until then the useful posture is neither dismissal nor enthusiasm. It is knowing which parts are load-bearing.
Key takeaways
- A July 2026 critical survey of 45 GEO studies found no technique with a stable, cross-platform causal effect on organic discoverability.
- Only topical relevance and position in the context window carry strong support, and position applies only after retrieval has happened.
- The famous 40% statistics finding came from a fixed five-document context; the survey grades the general claim as rejected.
- C-SEO Bench found 3 of 54 method-and-domain combinations significantly positive, and none on question answering.
- SAGEO Arena found body-only optimisation reduced retrieval and citation at every stage it measured.
- Whether AI citations produce clicks or revenue is unresolved, graded very low confidence on a single quasi-experiment.
Put this to work with Pantheraa: SEO & AI Search · Answer engine optimisation · Ranking in AI Overviews.
GEO evidence — questions, answered.
Parts of it do, under conditions. Content that is genuinely relevant to a query and sits early in the model’s context gets cited more, and both effects are graded high confidence in the July 2026 survey. What has not been shown is that any specific tactic reliably improves organic discoverability across engines and over time. The survey found no technique meeting that bar.
No. The best-supported recommendation is to produce relevant, verifiable, clearly structured and technically retrievable pages, then measure retrieval, citation and fidelity separately. That work has value regardless of how the engines evolve. What should stop is buying fixed recipes sold with percentage guarantees, since fixed recipes are graded as generalising poorly.
It traces to position-adjusted word count moving from 19.3 to 27.2 under one treatment, in a testbed where five documents were already supplied to the generator. Retrieval was held constant, so the study could not measure discoverability. The survey grades the general claim as rejected because a relative maximum on one metric is being presented as a broad effect.
Nobody has published evidence either way. The July 2026 survey does not treat schema markup or llms.txt as tested levers, so confident claims in either direction rest on nothing measured. Both are cheap to implement and unlikely to hurt. Treat them as reasonable housekeeping, not as an intervention with a known effect.
The survey cites seven to eight repetitions per cell as a reasonable starting point, validated against a pilot estimate of variance. Multiply that by your engines, query paraphrases and conditions and the number climbs quickly. Fewer than that and you are measuring drift, since engines change 9 to 28% of decisions on repeated runs even at temperature zero.
It is a poor idea. Visibility decomposes into retrieval, context exposure and citation, and a tactic can improve one while damaging another, which is exactly what SAGEO Arena found for body-only optimisation. A single score cannot tell you which stage is failing, so it cannot tell you what to change.
That the guarantee cannot be honestly given, and why. The link from AI citation to clicks or revenue rests on one quasi-experiment whose placebo test returned p=0.16. Offer instead a defined body of work, a stated measurement protocol, and honest reporting of what moved. Most clients respond well to a boundary drawn clearly.
Fast. Engines showed Jaccard similarity of 0.34 to 0.42 in cited sources across a 45-day window, and AI Overviews showed 18% page overlap over two months against 45% for organic results. A paper describing behaviour from a year ago is describing a system that has since changed, which is why the survey argues for versioned, continuously updated benchmarks.
Ready to replace guesswork with a growth engine?
Book a 30-minute strategy call. We’ll show you exactly where your funnel is leaking, before you spend a dollar.