Published
What GEO means, and what it meant first
Generative engine optimization (GEO) is the practice of changing a web page, or a body of content, so that a generative AI system — one that answers a question in prose and attaches citations, such as Google AI Overviews, Google AI Mode, ChatGPT search, Perplexity or Microsoft Copilot — becomes more likely to retrieve it, cite it, or repeat its claims. That is the definition the term was born with, and it is narrower than the one sold.
The phrase comes from an academic preprint posted on 16 November 2023 by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, published at KDD 2024 under DOI 10.1145/3637528.3671900. No earlier documented use of the exact phrase was found for this page, and the paper does not claim to have coined it. Coinage matters less than scope: the paper framed GEO as a narrow, testable optimization problem, while the commercial usage that took hold from mid-2025 describes a service. Most of the confusion in this subject follows from treating the two as one thing.
Engine optimization — naming a discipline after the machine rather than the reader it serves — has produced this failure mode before: confident advice about a system whose operators publish almost nothing about how it picks sources.
What the KDD 2024 paper actually tested
The experiment ran on GEO-bench, a benchmark of 10,000 queries in an 8k/1k/1k train, validation and test split. For each query the authors fixed the model's context at the top five sources fetched from Google, then rewrote one of the five with GPT-3.5-turbo using one of nine methods — keyword stuffing, unique words, easy-to-understand, authoritative, technical terms, fluency optimization, cite sources, statistics addition and quotation addition — and measured how much of the generated answer could be attributed to the rewritten document.
The retrieval set was held constant. The optimized document sat in the context window before the rewrite and after it. The paper measures what happens to a source once a generative engine has already selected it. It does not test, and nowhere claims to test, whether any of these rewrites make a page more likely to be selected at all.
That distinction carries the entire subject. A generative answer is assembled in stages: the system decides whether to search, expands the question into machine-generated sub-queries the user never typed, retrieves documents against them from a conventional search index, reranks and allocates a handful into the model's context window, then writes an answer and attaches citations. The early stages decide whether a page is in the room; the late stages decide what happens once it is there. GEO as published addresses only the late stages, and the early ones are what the operators actually document.
Where the 40 percent came from, and what it measures
Table 1 of the paper reports a baseline of 19.3 for unoptimized content and 27.2 for its strongest method, quotation addition. That is a relative gain of 40.9%, which the abstract renders as visibility boosted “by up to 40% in generative engine responses.” The metric is position-adjusted word count: the share of a generated answer's words attributable to your source, discounted by where in the answer they appear. It is not traffic, not clicks, not rank, and not a count of citations in a live product.
That sentence is quotable without its conditions, and in May 2025 an a16z essay turned the term into a venture category and created commercial demand for exactly one citable number. The same essay supplied the category's other repeated figure — more than $80bn — which is the size of the existing search engine optimization market, repurposed as GEO's addressable market by an investor in GEO tooling. Neither number measures GEO spending or GEO results.
Something vendor material almost never mentions: the paper also reports improvements of up to 37% for the same methods on Perplexity.ai, which it calls a commercially deployed generative engine. That is the stronger of the two results, because it ran against a real product rather than a prototype, and it is the one nobody quotes. A figure gets repeated because it is round and high, not because it is the best evidence in its own source.
Two further rows cut against the way the table is used. Keyword stuffing was the only method to score below baseline, at 17.7 against 19.3, so the founding document of GEO argues against a tactic still sold under its banner. And the systems tested were GPT-3.5-turbo in 2023 and 2024, plus Perplexity as it existed then; the paper's limitations section says the methods may need to adapt as generative engines evolve. Read the paper before quoting its number.
What happened when the methods were tested again
The first serious independent replication is C-SEO Bench (Puerto et al., arXiv 2506.11097, NeurIPS Datasets and Benchmarks 2025). It tested ten methods from this family — statistics, quotes, citations, authoritative phrasing, fluency and others — across two tasks and six domains, and concluded that “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected.” The count is what to remember. Out of 54 method-and-domain cases, three showed a statistically significant ranking improvement.
The same benchmark tested what happens when more than one publisher plays: “as we increase the number of C-SEO adopters, the overall gains decrease, depicting a congested and zero-sum nature of the problem.” That is a benchmark result, not a measurement of the live web, but it is the only multi-actor evidence there is: any edge these methods confer is an edge over publishers who have not adopted them yet.
A 2026 critical survey of 45 GEO studies by Martinez adds the finding that should give pause to anyone rewriting body copy for citability. Reporting the SAGEO Arena result, it describes body-only optimization cutting top-20 retrieval presence by roughly 9% and top-10 by roughly 16%, while improving citation conditional on retrieval. A rewrite can win the argument inside the context window and lose the seat that got it there. The benchmark itself is C-SEO Bench.
What the operators document, and what they refuse to
Google's documentation for AI features, last updated 10 December 2025, states: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” It documents eligibility instead — a page must be indexed and eligible to be shown with a snippet — and publishes only preview controls such as nosnippet, data-nosnippet and max-snippet, all of which only remove content. Danny Sullivan put it plainly at WordCamp US on 28 August 2025: “Good SEO is good GEO, or AEO, AIO, LLM SEO, or LMNOPO.”
Microsoft is the dissenting voice. In February 2026 Bing became the first operator to name GEO in its own webmaster guidelines, framing it around content eligibility for grounding and reference in AI responses and noting that GEO guarantees citations no more than SEO guarantees rankings. What it did next is more informative: the same update renamed its keyword-stuffing section to cover artificially engineered language and added a section on prompt injection and AI manipulation. Bing's answer to GEO tactics was to extend its spam policy, not to publish a ranking factor. That page is rendered in JavaScript and could not be read directly here; the wording is as reported by Search Engine Journal on 27 February 2026 and should be re-verified before quoting.
OpenAI documents access and nothing else: any public website can appear in ChatGPT search, and sites opted out of OAI-SearchBot will not be shown in its answers. Google has moved the other way — its spam policies, last updated 28 August 2026, treat attempts to manipulate generative AI responses in Google Search as spam.
The silence is the finding. As of August 2026 no major operator documents any content characteristic that raises citation probability, how sources are ranked for inclusion in an answer, how many enter the context window, whether structured data influences selection, or any GEO-specific markup or protocol. Bing naming the concept is the closest anyone has come, and naming a concept is not publishing a mechanism.
Tactics sold as GEO with nothing behind them
Each of these is widely recommended. Each fails on the evidence available in August 2026.
- Keyword stuffing for AI answers. The founding paper scored it at 17.7 against a 19.3 baseline — the only method to underperform doing nothing — and Bing's February 2026 update extended its keyword-stuffing prohibition to language engineered for AI citation.
- Adding statistics because statistics get cited. This is the paper's statistics addition method, which scored 25.2 in its own setting and did not survive C-SEO Bench. The paper reports no verification that the figures its language model inserted were accurate. Inventing them is unevidenced, sits inside Google's stated definition of spam, and fails the first time a reader checks one.
- Publishing llms.txt. John Mueller, 17 June 2025: “FWIW no AI system currently uses llms.txt.” Originality.ai tracking across more than three million sites, published 2 July 2026 with data to May 2026, found 97% of llms.txt files received zero AI-crawler requests in the month measured.
- Writing in a register aimed at models rather than readers. C-SEO Bench tested that family directly and found it largely ineffective and often harmful to ranking. Three significant results in 54 cases is what a null effect looks like once you test enough combinations.
How to test a GEO claim before believing it
Before the tactics, the instrument. Repeated identical queries return overlapping source sets only 0.34 to 0.42 of the time, URL-level similarity between Google and Gemini runs 0.11 to 0.18, and 57.8% of repeated ChatGPT queries skip web search entirely, all per the Martinez survey of July 2026. Reddit's share of ChatGPT Search citations fell from 3.83% to 0.52% across four days in August 2026, and the analysts who found it said plainly that the data shows when the shift occurred, not why.
A single before-and-after test against that noise floor establishes nothing. The minimum credible design is repeated measurement across paraphrased queries, over weeks, against an unmodified control, reporting a distribution rather than a position. Almost no published GEO case study meets it.
Three questions separate a claim worth reading from one worth ignoring. Does it measure retrieval or generation — did the page get into the context window, or merely do better once inside? Was anything left alone as a control? And how many times was each query run? Ahrefs shows why the last matters: its July 2025 analysis of 1.9 million citations put 76.10% of AI Overview-cited pages in the top 10, while a later one reported in March 2026 put the share at 38% — a gap Ahrefs attributes partly to changed parsing, so the two are not a trend line.
What would change this verdict: a study that starts from an unmodified live corpus, applies these rewrites, and measures citation rate in a commercial generative engine over time, with controls and repeated sampling. Nothing of that shape has been published.
Frequently asked questions
Does generative engine optimization actually work?
As a label for making content retrievable and worth citing, yes, because that is search engine optimization. As a distinct set of content rewrites, the evidence is against it. C-SEO Bench tested ten such methods across two tasks and six domains in 2025 and found statistically significant ranking improvement in three of 54 cases, with several methods actively hurting the document's ranking.
Where does the 40% GEO figure come from?
Table 1 of Aggarwal et al., posted to arXiv in November 2023 and published at KDD 2024. Quotation addition scored 27.2 against a 19.3 baseline on position-adjusted word count, a relative gain of 40.9%. That metric is the share of a generated answer attributable to your source in a fixed five-document context. It is not traffic, clicks, rank or citations in a live product.
Does Google recognize GEO as a discipline?
No. Google's AI features documentation, last updated 10 December 2025, states there are no additional requirements to appear in AI Overviews or AI Mode and no other special optimizations necessary. Danny Sullivan has said publicly that good SEO is good GEO. Microsoft's Bing is the only operator to name GEO in its guidelines, which it did in February 2026 without publishing a ranking factor.
Should I add statistics and quotations to my pages for AI citation?
Only where they belong in the writing. Those were the founding paper's two strongest methods in its own fixed-context setting, scoring 25.2 and 27.2 against a 19.3 baseline, and they did not survive independent replication. The paper never reports checking whether the inserted figures were true. Inventing them is unevidenced, is inside Google's stated definition of spam as of August 2026, and is checkable by any reader.
Why do GEO tactics seem to work when people test them?
Because these systems are unstable enough that a single test measures noise. Repeated identical queries return overlapping sources only 0.34 to 0.42 of the time, and 57.8% of repeated ChatGPT queries skip web search entirely. One prompt, one run and one screenshot is a single draw from a distribution. Any tactic will appear to work about half the time under that method.
Do GEO gains last once everyone adopts them?
The only multi-actor evidence says no. C-SEO Bench reports that as the number of adopters increases, overall gains decrease, describing the problem as congested and zero-sum. That is a benchmark finding rather than a measurement of the live web, but it means any advantage these methods confer is an advantage over publishers who have not adopted them yet, not a durable property of the page.