Published
Four different things are sold as one number
AI search measurement is the practice of establishing whether, how often and in what form a site or a brand appears inside answers produced by generative surfaces — Google's AI Overviews and AI Mode, ChatGPT, Perplexity, Microsoft Copilot, Gemini. It is not one measurement but four, sold under a single label:
- Citation — a link to your URL appears as a source beneath or beside the answer.
- Mention — your brand is named in the answer prose, whether or not anything links to you.
- Impression — a platform counts your link as having been shown to a user.
- Referral — a person clicks and arrives on your site.
Different instruments produce each one and they move independently: a page can be cited without the brand being named, a brand can be named with nothing linked, and a citation can register as an impression nobody clicked. Semrush's 2026 index separates the middle two explicitly, defining mentions as how often a company appears in an answer and citations as which pages the platforms use as evidence. Only the fourth is worth money on its own.
One property separates all of this from twenty years of rank tracking: the thing being measured is generated, not retrieved from a fixed index. Two people asking the same question at the same moment can receive different answers citing different sources. Every figure in this field is therefore a sample from a distribution rather than a reading of a state, and almost no vendor reports it that way.
Only four places a number can come from
Every product in this category rests on one or more of four data sources, and knowing which one produced a figure tells you most of what you need to know about it.
Platform-reported data comes from the operator, covers only Google and Microsoft, and is census-level for your own property rather than a sample — the only category that is not an estimate. Prompt monitoring sends a curated list of prompts to conversational assistants on a schedule and scores the answers; it is the only way to observe ChatGPT, Claude or Perplexity at all, and the only way to measure mention rather than citation, and it is a survey. Results-page scraping detects the AI Overview block inside a Google results page and extracts its citations, which is how presence rates across millions of keywords are produced. Server-side observation reads analytics referrers and crawler hits in logs.
No category covers the field, and each is blind in a predictable direction: platform reporting misses everything outside Google and Microsoft, scraping misses everything that does not render in a results page, prompt monitoring depends as much on prompt selection as on your site, logs measure input rather than output, and analytics measures the small surviving fraction of output. A program built on one of them has a specific hole in it, knowable in advance.
What the two platform reports actually contain
Google's generative AI performance report arrived in Search Console on 3 June 2026 and is narrower than its billing: impressions only, by pages, countries, dates and devices. No clicks column, no clickthrough rate, no average position, no query dimension, and no way to separate AI Overviews from AI Mode. Google describes it as a view onto data already inside the Web search type — a diagnostic lens rather than a scoreboard.
The clicks exist. Google documents that clicking an external link in an AI Overview counts as a click, and the same for AI Mode, but they are pooled into the Web total with no filter — so the largest AI surface on the web reports no isolable traffic figure anywhere, in Google's tooling or yours. AI Mode follow-ups are at least counted as new queries and do appear in the main query report, unlabeled and identifiable only by conversational shape.
Microsoft went further and earlier. Bing Webmaster Tools AI Performance, in public preview since 10 February 2026, reports citation counts across Copilot and Bing's AI summaries plus the grounding queries the system used to retrieve — the sub-query layer Google exposes nowhere. Citation Share followed on 16 June 2026: the percentage of citations attributed to your site out of all citations for that grounding query, and the only share-of-voice figure any platform publishes. Microsoft states plainly that the data does not indicate ranking, authority or page importance — worth repeating to anyone selling a score that weights citations by position in the answer.
One measurement is one draw
The variance has been measured, and it is worse than most reporting assumes. SE Ranking parsed the same 10,000 keywords in Google AI Mode three times on a single day in June 2025: 9.2% of cited URLs appeared in all three runs, any two runs shared roughly 18.5%, and domain-level agreement across all three was 14.7%. Ahrefs, across 540,000 query pairs in December 2025, found 45% of AI Overview citations change between generations of the same answer.
Part of the variation is now designed in. Since 27 May 2026 a user's Preferred Sources selections are reflected inside AI Overviews and AI Mode, making part of the citation set user-configured; Semrush's methodology notes that AI Mode may also tailor responses to search history and location. Two users are not supposed to agree. Query fan-out then breaks the unit of analysis: a citation decision is often made against a machine-generated sub-query nobody outside Google sees, so a tracker reporting that you are not cited for a phrase may be reporting an outcome decided on a different one.
And the surfaces are not proxies for one another: AI Overviews and AI Mode overlap at 13.7% of citations by Ahrefs' measure and 10.7% at URL level by SE Ranking's. The consequence is one sentence. A week-on-week change reported without runs-per-prompt and an error bar is variance presented as signal.
Where organic CTR goes wrong in these reports
Organic clickthrough rate is organic clicks divided by organic impressions. It is the metric most often reached for here, and it breaks for three separate reasons.
First, the counting rules differ. An AI Overview impression registers in Search Console only when the link is scrolled or expanded into view, while a classic organic impression is counted when the result set is generated. The two are not on the same scale and no ratio between them means anything. Second, generative AI impressions are a subset already inside the Web total rather than a parallel channel, so dividing one by the other to produce an “AI share” is arithmetic without a referent. Third: the report has impressions and no clicks, so an AI Overview clickthrough rate cannot be computed from it at all. Any dashboard showing one is showing an inference, and the method should be demanded and read.
The wider trap is treating a falling ratio as lost traffic. Seer Interactive reported organic clickthrough on AI Overview queries falling from 1.76% to 0.61% in November 2025, and that 61% decline is the figure that traveled; queries without an AI Overview fell 41% in the same window, and that figure did not. Seer's fourth-quarter data then showed impressions rising from 15.8 million to 33.1 million while clicks moved from 398,798 to 400,271. Publishers are paid in the numerator.
The attribution floor
Referral measurement is the only part of this that connects to revenue, and it is a floor rather than a measurement. The largest surface is the least attributable: clicks from AI Overviews and AI Mode arrive with a google.com referrer, indistinguishable from organic traffic in any analytics tool. For most sites that is the majority of AI exposure, and it is structurally invisible.
Exactly one operator has made an attribution commitment: OpenAI documents that ChatGPT includes utm_source=chatgpt.com in referral URLs, which makes it the best-instrumented surface by a distance. Every other hostname in an AI segment — Perplexity, Copilot, the Gemini app, Claude — is reverse-engineered convention no operator has promised, and Microsoft has never documented what referrer Copilot sends in more than two years of it sending one.
Then there is the leakage. Native apps frequently send no referrer, referrer policy can strip the header, people copy URLs out of answers, and an answer that names a brand without linking it produces a visit with no referral event to lose. That last case is not broken tracking; it is the outcome these products are designed to produce, and no instrument can see it. The one quantified estimate of the gap rests on a single consultant's unpublished method, so report the mechanisms confidently and the magnitude as unmeasured.
For scale: SE Ranking's 101,574-site dataset put AI referrals at 0.02% of website traffic in 2024 and 0.32% in 2026 — sixteen-fold growth on a base of a third of one percent, and quoting either half alone misleads in opposite directions. The same data shows 60% of AI-referred traffic landing on homepages against 17% from organic, so page-level analysis will tell you little about which page earned the recommendation.
Logs measure the other side, and their limits are not only technical
Server and CDN logs are the only census of the input side: what the AI systems actually fetched. OpenAI publishes four user agents with distinct meanings and an IP range file for each, so identification can be verified rather than assumed. GPTBot is training; OAI-SearchBot builds the index behind ChatGPT search; ChatGPT-User is a live, user-initiated fetch of your URL — evidence your page was pulled into a real conversation, though not proof of citation or of a visit. Counting GPTBot hits as visibility confuses training harvest with answer-time retrieval. Logs also separate a retrieval failure from a selection failure, which nothing else does.
Three limits. Nobody has published a ratio between crawling and citation. No Google user agent isolates AI crawling: Googlebot covers all Search features including AI Overviews and AI Mode, while Google-Extended governs training and grounding for Gemini Apps and Vertex AI, Google's enterprise platform for building and deploying models, and does not affect inclusion in Search. Any report showing “Google's AI crawler” is showing Googlebot.
Third, a user-agent count is a lower bound rather than a census, and the evidence is now in court. Cloudflare attributed 3 to 6 million daily requests to undeclared Perplexity crawlers in August 2025, against 20 to 25 million from the declared one; Perplexity disputed it. News Corp's Dow Jones and New York Post had sued Perplexity in October 2024, and Nikkei and Asahi Shimbun filed in Japan in August 2025 alleging robots.txt was ignored. Where crawl conduct is contested in litigation, treating header counts as complete is untenable.
What a defensible report discloses
There is no standardized metric here and no platform defines one. Vendors define their own from different inputs, over different prompt sets, at different sampling rates. Semrush's own index moved from 2,500 prompts to 126 million between versions — a denominator that changes by five orders of magnitude does not produce a time series, and two vendors reporting share of voice for the same brand in the same month can differ by a large multiple with neither being wrong.
A report worth acting on says on its face what it measured and how:
- Which of the four — citation, mention, impression or referral — it is, and the surface and date it came from.
- The prompt set and how it was chosen, the runs per prompt, and how they were aggregated. A tool that will not say how many times it sampled is reporting a coin flip.
- Geography, device and account state, and whether the consumer product or an API was queried — different systems with different retrieval.
- The floor, marked as a floor. Referral figures exclude Google's own AI surfaces and lose app traffic, stripped referrers and every mention that was never linked.
Two first-party inputs are free and routinely skipped: Search Console's impressions and Bing's citation and grounding-query data, both census-level for your own property. Everything else samples a black box. The honest version of this work is not a single visibility score but a small set of clearly labeled measurements, each with its blind spot written beside it.
Frequently asked questions
Can I see AI Overview clicks in Search Console?
No. Google documents that clicks from AI Overviews and AI Mode are counted, and pools them into the Web search type with no filter. The generative AI performance report added in June 2026 shows impressions only, with no clicks, no clickthrough rate, no average position and no query dimension. The data exists inside Google and is not exposed in a form anyone can isolate.
Why do two AI visibility tools give my brand completely different scores?
Because they measured different prompt sets a different number of times, possibly on different surfaces, and defined appearance differently. Both can be internally consistent and still disagree by a large multiple. The comparison is meaningless without both methodologies, and most are not published — which is why the disclosure list matters more than the score.
How many times does a prompt need to be checked?
More than once, and how many more depends on what you can tolerate. SE Ranking found only 9.2% of cited URLs common to three parses of the same 10,000 keywords on the same day in AI Mode, and Ahrefs found 45% of AI Overview citations changing between generations of the same answer. A single pull is one draw from a distribution whose variance was never reported.
Is AI referral traffic worth tracking when it is under one percent?
Yes, provided both halves of the number travel together. SE Ranking's 101,574-site dataset put AI referrals at 0.32% of website traffic in 2026, up sixteen-fold from 0.02% in 2024. The base is small and the trend is steep, and the figure is a floor in any case, because it cannot see Google's own AI surfaces or any recommendation that named a brand without linking it.
Does blocking AI crawlers show up as a drop in visibility?
It shows up in different places depending on what was blocked, and one common configuration does nothing at all. Google-Extended covers training and grounding for Gemini Apps and Vertex AI, and Google says it does not affect inclusion in Search, so it will not remove a site from AI Overviews. Blocking OAI-SearchBot does remove a site from ChatGPT search answers, by OpenAI's own statement. Blocking ChatGPT-User mainly destroys the one live-retrieval signal available in logs.
Is there a single metric for AI search visibility?
No platform defines one and no standard exists. Microsoft's Citation Share is the only share-of-voice figure any operator publishes, and it covers Copilot and Bing only, reports citations rather than traffic, and exposes no competitor names. Any composite score is a vendor's own construction, and it stops being comparable to itself the moment the vendor changes its prompt set.