A technical reference on AI search visibility. This site sells nothing, takes no engagements and endorses no products. Consulting enquiries are handled separately at hartzer.com.

Hartzer.it.com logoHartzer.it.comAI search visibility reference
Abstract funnel form illustration representing Optimizing for AI search
Pillar guide

Optimizing for AI search

Optimizing for AI search is search engine optimization plus a few access decisions. This is what the operators document, and what replication did to the rest.

DocumentedGoogle documents what is required to appear, and the documentation says there is nothing beyond ordinary search engine optimization.

Start with the only requirement anyone has published

Google's guide to optimizing for generative AI features carries a section headed Mythbusting generative AI search: what you don't need to do — an explicit list of the things Google says are unnecessary. The first entry is structured data: “Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add.” Its companion page, last updated 10 December 2025, states that “there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary,” and that to be eligible as a supporting link a page “must be indexed and eligible to be shown in Google Search with a snippet.”

That is the entire documented requirement, and it is a restatement of search engine optimization — the practice of making pages discoverable, crawlable and rankable in a search engine's ordinary index. Optimizing for AI search, as any operator has described it, is search engine optimization plus a small number of access decisions. Everything sold on top of that is a hypothesis, and the useful question is which hypotheses have survived a test.

The operators are consistent on this in public as well as in documentation. Danny Sullivan, at WordCamp US on 28 August 2025: “Good SEO is good GEO, or AEO, AIO, LLM SEO, or LMNOPO,” and on Search Off the Record in December 2025 he placed the newer discipline inside the older one. Microsoft, which on 27 February 2026 became the first operator to name generative engine optimization in its own webmaster guidelines, did so while noting that it no more guarantees citations than search engine optimization guarantees rankings.

Where the vocabulary came from

Generative engine optimization, or GEO, is the practice of changing a page so that a system which answers questions in prose is more likely to retrieve it, cite it, or repeat its claims. The term is academic in origin: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande posted it to arXiv on 16 November 2023, and the paper was published at KDD 2024. Answer engine optimization, or AEO, is older and murkier — the earliest verifiable dated use is Rebecca Sentance in Search Engine Watch on 7 February 2018, and she wrote in the passive, “this strategy has come to be known as AEO,” which means it was already circulating and she was not the coiner. The frequently repeated attribution to a 2017 coinage has no traceable support.

The common element is the phrase engine optimization, borrowed grammar from a twenty-year-old discipline. Each new label keeps those two words and swaps the adjective, which is a positioning move rather than a mechanical distinction, and it is worth noticing who benefits from each swap. The commercial category did not begin with the paper. It began on 28 May 2025, when a16z published an essay recasting GEO as a market, and the widely quoted addressable figure in it is the size of the existing search engine optimization market repurposed as GEO's opportunity by an investor in the category. No operator has ever named AEO in documentation at all.

The pipeline has two places to win, and most advice targets one

A generative answer is produced in stages. The system decides whether to search at all; the question is expanded into sub-queries; documents are retrieved against those; a subset is reranked and placed in the model's context window; the model writes; and citations are attached. The first four determine whether your page is in the room. The last two determine what happens once it is there. Nearly every argument in this subject collapses when those two loci are separated.

The founding GEO paper measured only the second locus. It held the retrieval set fixed at the top five documents fetched from Google, rewrote one of them with GPT-3.5-turbo, and measured citation within that fixed set. The famous “up to 40%” is 27.2 against a 19.3 baseline on position-adjusted word count — the share of the generated answer's words attributed to your source, discounted by where in the answer they appear. It is not traffic, not clicks, not rank, and not citation count in a commercial product. Your document was already retrieved before the experiment started, and the paper neither tests nor claims to test whether any rewrite makes retrieval more likely.

That distinction is not academic hygiene. It has a measured cost. Martinez, reporting the SAGEO Arena result in a 2026 survey, found that body-only optimization reduced top-20 retrieval presence by roughly 9% and top-10 presence by roughly 16%, while improving citation conditional on having been retrieved. A rewrite can win the last step of the pipeline and lose an earlier one, and the earlier one is the step that decides whether you exist in the answer at all.

What replication did to the tactic list

The tactic inventory in circulation — add statistics, add quotations, add citations, write authoritatively, improve fluency — comes almost entirely from the nine methods in that one paper. It has since been tested properly, and it did not hold up.

C-SEO Bench (Puerto, Gubri, Green, Oh and Yun, NeurIPS Datasets and Benchmarks 2025) ran ten of those methods across two tasks and three domains each, and added the design contribution nobody else had made: varying how many competitors adopt the tactic. The conclusion is unusually blunt for a benchmark paper — “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected.” Out of 54 method-and-domain cases, three showed statistically significant ranking improvement. And the gains that did appear shrank as adoption rose, which the authors describe as a congested, zero-sum problem.

Then the systematic review. Martinez surveyed 45 GEO studies and published on 15 July 2026, concluding that already-retrieved content can causally alter its own citation, but that “no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.” Note also what the founding paper itself found and its resellers do not repeat: keyword stuffing was the only one of the nine methods to score below baseline, at 17.7 against 19.3. Bing extended its keyword-stuffing prohibition to cover “artificially engineered language” aimed at AI citations in February 2026.

Three things sold as AI optimization with nothing behind them

Structured data for AI citation. Ahrefs ran the only controlled test on 11 May 2026: 1,885 pages that added JSON-LD between August 2025 and March 2026 against roughly 4,000 matched control pages, measured 30 days either side. The result was AI Overviews −4.6%, AI Mode +2.4%, ChatGPT +2.2% — no uplift on any platform. The same study also found that cited pages were about three times as likely to carry JSON-LD, and that 53% of AI-cited pages carried it, which is the half of the study that gets quoted. There is a mechanism argument too: systems that fetch a live URL and convert it to text typically discard <script> blocks, which is where JSON-LD lives. Google's own optimization guide says structured data is not required for generative AI search — and in the next sentence says keep using it, because it still earns rich results.

llms.txt. Gary Illyes said in July 2025 that Google “doesn't support LLMs.txt and isn't planning to”; John Mueller wrote in June 2025 that no AI system currently uses it; Originality.ai's tracking across more than three million sites found 97% of llms.txt files received zero AI-crawler requests in May 2026. It spread because it is cheap, it looks like robots.txt, and shipping one is legible as work. What would settle it is server logs showing an AI crawler requesting the file. Nobody has produced them.

Statistics added for the sake of having statistics. The founding paper's strongest methods worked by adding quantitative material through a language model, and the paper reports no verification that the added figures were accurate. Google's spam policies, updated 28 August 2026, treat attempting to manipulate generative AI responses as spam. A fabricated number is a policy exposure and a reputational one the first time a reader checks it.

What is left, and it is not nothing

Strip out the unsupported layer and a short list survives, most of it unglamorous.

  • Eligibility. Indexed, crawlable, snippet-eligible. Google states it as the precondition, which means an indexation problem is an AI visibility problem and is usually the whole of one.
  • Access decisions, made deliberately. OpenAI states that sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers.” Cloudflare began blocking AI crawlers by default for new domains on 1 July 2025, so a number of sites are opted out of surfaces they believe they are competing in. And the control most people reach for is the wrong one: Google-Extended governs Gemini Apps and Vertex AI training, not AI Overviews, and Google's documentation still says it does not affect inclusion in Search. The actual opt-out arrived in Search Console on 17 June 2026.
  • Ranking, because ranking is retrieval. C-SEO Bench found that getting a document to the front of the model's context beat every rewrite tested — and getting to the front of the context is a classical ranking outcome, not a writing one.
  • Unambiguous prose. Microsoft is the only operator to give content instructions for grounding, and they are modest: state facts directly rather than implying them, and keep entity names clear and consistent with no ambiguous references.
  • Structured data, for what it actually does. Rich results still exist and still pay. Google says to keep the markup, and that unused markup causes no problems.

The cost of getting this wrong

The usual defense of AI-specific tactics is that they are cheap and might work. Both halves are weaker than they look.

They are not free of downside. Body-only rewriting measurably cost retrieval presence in the one study that separated the two effects. Whatever gains exist compress as adoption rises, so a tactic is worth most in the window before it is widely sold and least once everyone has bought it. And the budget spent on markup with no live consumer, or on a file no crawler requests, is budget not spent on the indexation and ranking work the documentation says governs eligibility.

They are also, in most cases, untestable as deployed. Martinez reports source-level overlap of 0.34 to 0.42 across repeated identical queries, and that 57.8% of repeated ChatGPT queries skip web search entirely. Against that noise floor, a single before-and-after measurement of a page change cannot distinguish an effect from ordinary variance. Anyone reporting that a rewrite “worked” on the strength of one check is reporting a coin flip with a narrative attached, and the same applies to anyone reporting that it failed.

What would change the verdict is specific, and nobody has published it: a study that starts from an unmodified live corpus, applies the rewrites, and measures citation rate in a commercial generative engine over time, with a control group and repeated sampling. Until then the honest position is the one the operators have already written down — do the ordinary work well, make the access decisions on purpose, and treat the rest as an experiment you are paying to run.

Frequently asked questions

Is GEO a different discipline from SEO?

Google's public position is that it is not: Danny Sullivan has called good SEO good GEO and has placed GEO inside SEO as a subset. Microsoft names generative engine optimization in its guidelines but frames it as content eligibility for grounding, and says it guarantees citations no more than SEO guarantees rankings. The mechanics support them — the documented requirement for appearing in Google's AI features is being indexed and eligible in ordinary Search.

Does adding schema markup get a page cited more often?

Not in the one controlled test that exists. Ahrefs measured 1,885 pages that added JSON-LD against roughly 4,000 matched controls in May 2026 and found AI Overviews down 4.6%, AI Mode up 2.4% and ChatGPT up 2.2% — no meaningful uplift anywhere. Cited pages are about three times more likely to carry schema, but that is a correlation between two consequences of being a well-maintained site. Keep markup for rich results, which is what Google recommends.

What about the study showing GEO increases visibility by 40%?

It measured something narrower than the number suggests. The retrieval set was fixed at five documents, one of them was rewritten, and the metric was position-adjusted word count — the share of the answer's words attributed to that source. The document was already retrieved. The paper does not test whether any rewrite improves the odds of being retrieved, which is the step that decides whether you appear at all.

Should I publish an llms.txt file?

There is no evidence it does anything. Google has said it does not support the file and is not planning to, and tracking across more than three million sites found 97% of llms.txt files received zero AI-crawler requests in May 2026. It costs little, which is why it spread, but it also does nothing that has been demonstrated, and it is not a substitute for making sure crawlers can reach the pages themselves.

Which crawler control actually removes a site from AI Overviews?

The search generative AI control in Search Console, effective 17 June 2026. Not Google-Extended, which was introduced in September 2023 for Gemini Apps and Vertex AI training and which Google's documentation still says does not affect inclusion in Google Search. Setting Google-Extended and believing you have opted out of AI Overviews is one of the most common configuration errors in this subject.

If most tactics do not work, what should the effort go into?

The documented layer: being indexed and snippet-eligible, being reachable by the crawlers whose surfaces you care about, ranking well enough to be retrieved in the first place, and writing prose that states facts directly and names entities consistently. That is a shorter list than most AI-visibility proposals contain, and it is the part with either operator documentation or benchmark evidence behind it.

Top