Published
Three decisions hidden inside one citation
Source selection in an AI answer is a problem in information retrieval — the discipline of finding documents that satisfy a query — with an extra step bolted on the end. A generative surface decides which pages to retrieve, which of those to actually use when composing prose, and which to show the reader as a citation. Those are three decisions, not one, and they do not produce the same set of pages. A page can be retrieved and never used, used and never cited, or cited while contributing almost nothing to the sentence it sits beside.
One term does most of the damage here. A query is the string a retrieval system actually runs, and since 2025 that is frequently not the string the user typed. This is why publisher diagnosis so often goes wrong: a page that was never fetched and a page that was fetched but not credited look identical from the outside, and they call for opposite responses.
Every major surface — Google AI Overviews, Google AI Mode, ChatGPT search, Perplexity, Microsoft Copilot — performs all three steps, and none publishes the ranking function governing any of them. What they publish is an eligibility rule and a disclaimer. The useful question is therefore not "what is the algorithm" but "how much of the classic organic ranking system carries over" — and that has been measured, repeatedly, by parties who disagree.
Eligibility: the one rule Google publishes
Google states the precondition directly in its Search Central documentation, on a page whose last-updated stamp read 10 December 2025 when checked in August 2026: to be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There is no separate AI index, no separate submission route, no separate opt-in, and — Google says this in the same document — no special schema.org structured data and no new machine-readable file.
Microsoft made the same structural move in February 2026, rewriting the Bing Webmaster Guidelines to cover Copilot and grounding results inside the existing document rather than publishing a parallel one. Two operators, two years apart, both declining to create a second discipline.
This is the most useful documented fact in the subject and it is routinely skipped because it is boring. It collapses a large volume of AI-specific advice back into ordinary indexation work: if a section of a site is not indexed, is blocked, or is snippet-suppressed, no amount of answer-shaped rewriting will put it in a generated response, because it never reaches the stage where any of that could matter.
You are not competing for the query the user typed
Query fan-out — the practice of decomposing one question into several sub-queries, running them concurrently, and composing a single answer from the combined results — is Google's own term. It first appeared publicly on 5 March 2025 in the AI Mode announcement and entered Google's Search Central documentation in December 2025, hedged as something both AI Overviews and AI Mode "may use". OpenAI documents the same operation in different words: ChatGPT "rewrites your query into one or more targeted queries" before sending them to its search providers.
The consequence for source selection is direct. The candidate pool is assembled per sub-query, not per original query. A page that could never rank for the question the user asked can be pulled in because it ranks for a question the system invented and nobody sees.
Two cautions, because this is where the subject fills with invention. Google has given exactly two figures: Deep Search, a separate and slower feature, "can issue hundreds of searches" (May 2025), and an engineering director said in an interview post in March 2026 that AI Mode is "basically doing a dozen searches for you in the time it takes to do one". Every more precise number in circulation — eight variant types, twenty iterations, one query becoming twelve — comes from practitioner readings of Google patents, not from Google. And the sub-queries themselves are visible nowhere: Search Console surfaces AI Mode follow-up questions typed by users, not machine-generated sub-queries. Tools sold as revealing fan-out are prompting a language model for plausible sub-questions. That may be useful topic planning. It is not a measurement of Google.
How much of classic ranking carries over: four studies, four answers
This is the one part of source selection that outsiders can measure, by scraping answers and joining the cited URLs to rankings. Four measurements, all with published samples, all pointing different distances in different directions:
- Ahrefs, 2 March 2026 — 863,000 keyword result pages and roughly four million AI Overview URLs. 37.9% of citations came from URLs in the first ten results, about 31% from positions 11 to 100, and about 31% from URLs ranking nowhere in the top 100. The same study run in July 2025 put top-ten overlap near 76%.
- seoClarity, October 2025 — 362,000 US desktop queries. 32% overlap with the top ten, but 94% of overviews carried at least one citation from the top twenty, and position-one URLs appeared 43% of the time, decaying to 7% at position twenty.
- BrightEdge, 18 September 2025 — a series moving the opposite way: 32.3% in May 2024 rising to 54.5% in September 2025, with wide industry spread (healthcare 75.3%, restaurants 19.2%). Sample size not disclosed.
- Semrush, 21 July 2025 — 5,000 keywords and more than 150,000 citations, measuring domain and URL overlap separately: AI Mode roughly 54% domain against 35% URL; AI Overviews 86% against 67%; Perplexity 91% against 82%.
Two judgements follow, and neither is in the headlines these studies generated. First, the Ahrefs pair is not a trend line. Ahrefs said its parsing methodology changed between the two runs, so 76% to 38% is two numbers from two parsers, not a decline you can date. Anyone presenting it as a clean collapse has dropped the authors' own caveat.
Second, the Semrush split is the most operationally useful of the four and gets the least attention, because it separates two different failures. High domain overlap with low URL overlap means the engines agree with the rankings about which sites are credible and disagree about which page answers the question. That is a content-architecture problem, not an authority problem, and the two are treated as the same thing constantly.
From retrieved to credited: the step nobody documents
Citation is a display decision made after the answer is written, and it is the least documented of the four steps. Perplexity — the product whose entire value proposition is attribution — describes its source selection in two sentences: it searches the internet, gathering information from "authoritative sources", and attaches numbered citations. It does not define authoritative, does not say whose index it searches, and has published no ranking documentation. Repeated searching in August 2026 found nothing further. The silence is the finding.
Google says "core information quality systems" and stops. Semrush measured the display side in July 2025: 92% of AI Mode responses showed sidebar links covering about seven domains, 7% carried links below the answer, 1.7% carried none. Seer Interactive named the gap in March 2026 — a ghost citation, where content is retrieved and relied on but the brand is never named. If that is your actual failure, optimizing for citation is optimizing the wrong step, and no available tool will tell you which of the two is happening.
There is also a limit on what a citation means once you get one. Liu, Zhang and Liang, in Findings of EMNLP 2023, hand-evaluated four generative search engines and found that "a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence". That study covered Bing Chat, NeevaAI, perplexity.ai and YouChat as they existed in 2023 — not AI Overviews, not AI Mode — and the figure is now quoted as though it described Google's current behavior. No equivalent audit of today's surfaces has been published, which is worth saying plainly. The general point stands untested rather than fixed: a citation is a pointer, not an endorsement.
What the tested tactics actually do
Three pieces of controlled work exist, and all three cut against the advice built on top of them.
C-SEO Bench (Puerto, Gubri, Green, Oh and Yun, NeurIPS Datasets and Benchmarks 2025) tested ten content-rewriting methods across two search tasks and three domains each, and added the design element nobody else has: an evaluation protocol that varies how many competitors adopt the same tactic. The findings are blunt. Most methods "are not only largely ineffective but also frequently have a negative impact on document ranking". Three of 54 method-and-domain cases reached statistical significance. Gains shrink as adoption rises. And getting a document to the front of the model's context window produced far larger citation gains than any rewrite tested — which is a retrieval outcome, not a writing outcome.
The founding 2023 GEO paper's "up to 40%" is the most misused number in the subject: a relative improvement on a position-adjusted word-count metric, for a document already sitting in a fixed five-document context, against a research prototype. It says nothing about whether a rewrite makes a page more likely to be retrieved, and the paper does not claim otherwise.
On structured data, Ahrefs ran the closest thing to a controlled test in May 2026: 1,885 pages that added JSON-LD between August 2025 and March 2026, against 4,000 matched controls. AI Overviews fell 4.6%, AI Mode rose 2.4%, ChatGPT rose 2.2% — no meaningful uplift on any platform. That result is consistent with Google's own documentation, which says in writing that no special structured data is needed. Anyone selling a structured-data program as an AI citation service is selling something the platform has denied is required and a controlled test found no effect for. The authors' own limitation is worth keeping: their sample was pages already heavily cited, so the test does not address whether markup helps an invisible page become visible.
What would let anyone answer this properly
Only one operator publishes per-page citation counts. Microsoft's AI Performance report in Bing Webmaster Tools, in public preview since 10 February 2026, reports total citations and grounding queries across Copilot and Bing's AI summaries, and Microsoft states the data reflects citation patterns rather than ranking or page importance — a caveat worth borrowing. Google's equivalent, live 3 June 2026, gives impressions only, merges AI Overviews with AI Mode, and carries no query dimension and no clicks.
Two limits govern how much weight third-party citation research can carry. The systems are unstable — SE Ranking parsed the same 10,000 AI Mode results three times in one day in June 2025 and found 9.2% of cited URLs common to all three pulls — and since 27 May 2026 Google's Preferred Sources feature makes part of the citation set user-configured, so two people are not looking at the same answer by design.
What would settle the open question — what selects a source once retrieval has happened — is an operator publishing it, or an independent party isolating it with repeated sampling and controls. Neither has happened. The honest position is that eligibility is documented, decomposition is documented, and the step in between is a black box that several vendors will confidently describe to you.
Frequently asked questions
Do I need special markup to be cited in AI Overviews?
No. Google's documentation states that there is no special schema.org structured data you need to add and no new machine-readable file required. A controlled Ahrefs test in May 2026, comparing 1,885 pages that added JSON-LD against 4,000 matched controls, found no meaningful citation uplift on any platform.
Does ranking first guarantee a citation?
No. seoClarity's October 2025 data on 362,000 US desktop queries found position-one URLs appearing in an AI Overview 43% of the time, falling to 7% at position twenty. High rank raises the probability substantially and settles nothing.
Why do AI answers cite pages that do not rank for the query?
Because the retrieval runs against sub-queries the system generated, not against what the user typed. Google documents this as query fan-out. A page that ranks well for a related sub-question can enter an answer for a question it would never rank for.
Can I see which sub-queries Google ran for my page?
No. No consumer surface exposes them, and Search Console shows user follow-up questions inside AI Mode rather than machine-generated sub-queries. The Gemini API does return the queries a model executed, but that is a developer surface and cannot be used to audit a consumer result.
Which citation-overlap figure should I use?
Whichever one you can name, date and describe the sample for. The published measurements run from 32% to 54.5% and were taken on different dates, in different countries, on different query mixes, with proprietary parsers. A bare percentage in this subject always has a matching contradiction somewhere.
Does being cited mean the engine agrees with what my page says?
Not reliably. The one human evaluation of citation accuracy, published in Findings of EMNLP 2023, found 51.5% of generated sentences fully supported by their citations and 74.5% of citations supporting the sentence they were attached to. That covered four 2023 products, and no equivalent audit of the current surfaces has been published.