Published
Four different measurements sold under one name
AI visibility tracking is the practice of measuring whether, how often and in what form a brand or a website appears inside answers produced by generative search surfaces — Google AI Overviews and AI Mode, ChatGPT, Perplexity, Microsoft Copilot and Gemini among them. It is a measurement discipline rather than a product, and it exists because the answer a person reads is assembled at request time rather than looked up in a fixed index.
The term covers four quantities that are produced by different instruments and move independently of one another:
- Citation — a link to your URL appears as a source in the answer.
- Mention — your brand name appears in the answer prose, whether or not anything links to you.
- Impression — a platform counts your link as having been shown to a user.
- Referral — a person clicks through and arrives on your site.
A page can be cited without the brand being named, a brand can be named without being cited, and a citation can be counted as an impression that is never clicked. Semrush's June 2026 AI Visibility Index keeps two of them apart on purpose, defining mentions as "how often a company appears in an answer" and citations as "which domains and pages AI platforms use as evidence". Conflating the four is the origin of most of the bad numbers in the category.
What separates this from twenty years of rank tracking is that the object being measured is generated rather than retrieved. Two people issuing the same query at the same moment can receive different answers citing different sources, so every figure here is a sample drawn from a distribution rather than a reading of a state. Almost no vendor reports it that way.
Where any number you are shown actually came from
There are four places a measurement in this field can originate, and knowing which one produced a figure tells you most of what you need to know about how far to trust it.
- Platform-reported data. The operator tells you directly. Only Google and Microsoft do this at all, and each reports a different fragment.
- Prompt monitoring. A tool sends a list of prompts to each surface on a schedule and records whether the brand was named and the domain cited. This is a survey — a measurement taken on a sample the measurer chose and generalized to a population nobody has enumerated — and the output depends on which prompts were selected, how many times each ran, from where, and under what account state.
- Results-page scraping. AI Overviews render inside a Google results page, so a tracker that already scrapes results pages can detect the block and extract its citations. It reaches AI Overviews, partially reaches AI Mode, and cannot reach ChatGPT, Claude or Perplexity, which have no scrapeable page.
- Server-side observation. Referrals arriving at your site and crawler hits arriving at your origin. Referrals measure outcomes and undercount badly; crawler hits measure input and prove nothing about output.
Generative engine optimization, or GEO — adapting content in the hope of being selected by an answer engine — is judged almost entirely on numbers from the middle two sources. Both are samples, and neither is a census of anything.
What Google reports, and what it counts but will not show you
Since 3 June 2026 Search Console has carried a generative AI performance report for Search. Google's definition of its single metric is exact: "Impressions are how many times links to your site were shown to a user in a generative AI feature on Google Search." That is the whole report. It has no clicks, no click-through rate, no queries and no average position, and four documented dimensions — pages, countries, dates and devices (Search Console Help).
Two consequences are regularly missed. The report carries no dimension separating AI Overviews from AI Mode, so any figure taken from it covers both surfaces together. And it is a view onto data already inside the Web search type rather than a separate collection, which makes the popular exercise of dividing generative AI impressions by Web impressions to produce an "AI share" a division of a subset by the total that already contains it. Rollout was also incomplete: John Mueller's comment on 11 August 2026 was "Correct, it's not yet every domain."
The clicks are real and Google counts them. The counting documentation states that "Clicking a link to an external page in the AI Overview counts as a click", and says the same of AI Mode. They are then pooled into the Web search type with no filter. Organic click-through rate, the ratio every search team has computed for two decades, is therefore not computable for generative surfaces from Google's own data: the numerator exists, is counted, and is not exposed.
One further sentence in the same documentation constrains everything built on top of it: "An AI Overview occupies a single position in search results, and all links in the AI Overview are assigned that same position." Google recognizes no ordering of sources inside an answer, which removes the foundation from any product reporting where your citation sits within a block.
What Microsoft reports that Google does not
Bing Webmaster Tools has carried an AI Performance report since 10 February 2026. It shows "the total number of citations that are displayed as sources in AI-generated answers" across Microsoft Copilot, AI-generated summaries in Bing and select partner integrations. It also exposes the grounding queries — the retrieval phrases the AI system actually issued, as distinct from the words a person typed. Google exposes that layer nowhere.
On 16 June 2026 Microsoft added Citation Share, defined as "the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query" (Bing Webmaster Blog). It is the only share metric in this field published by the party that owns the denominator.
Microsoft is also unusually direct about what the data is not. Citation counts "do not indicate ranking, authority, or the role of any page within an individual answer". Any visibility score that weights citations by where they appear in an answer is constructing an ordering the platform has said is not meaningful.
The asymmetry is worth sitting with: the smaller surface carries the better instrument. Microsoft reports citations, the retrieval queries behind them and a share figure with a real denominator, while Google reports impressions and nothing else. Neither reports the other's metric, so the two cannot be combined into one series without inventing the conversion.
Non-determinism is measured, not theoretical
The strongest reason to distrust a single reading is not an argument about how language models work. It has been measured at scale, twice, by two different companies.
SE Ranking parsed the same 10,000 keywords three times on the same day in Google AI Mode and found that 9.2% of URLs appeared in all three runs. Domain-level overlap across all three was 14.7%, and any two runs overlapped at roughly 18.5–19% on URLs (SE Ranking, 29 August 2025). Ahrefs, working from 540,000 query pairs of September 2025 US data, reported on 15 December 2025 that 45% of AI Overview citations change between generations.
Read together, those figures settle a methodological question most dashboards leave open. A single measurement of your citation status in AI Mode is a sample of size one drawn from a distribution in which roughly nine of every ten observed URLs will not recur on the next draw. Any week-on-week movement below a very wide margin is indistinguishable from regeneration noise unless the tool sampled repeatedly and states how many times it did.
Three further sources of variance compound it. Since 27 May 2026 a user's Preferred Sources selections are reflected inside AI Overviews and AI Mode, making part of the citation set user-configured. Query fan-out — the decomposition of one question into several sub-queries before an answer is written — means the citation decision is often made on a query the site owner never sees. And the surfaces are not proxies for each other: AI Overviews and AI Mode share 13.7% of citations by Ahrefs' measurement and 10.7% at URL level by SE Ranking's.
Where the trade press and the documentation disagree
On 4 August 2026 Search Engine Journal's explainer of the generative AI report stated that it "separates AI Overviews from AI Mode, treating them as distinct features". Google's help page for the same report lists four dimensions — pages, countries, dates and devices — and no surface dimension of any kind.
That disagreement is recorded here rather than resolved, because the documentation is the only thing either party has published. What can be stated precisely: as of 29 August 2026 Google's own reference describes no way to split the report by surface, and anyone who needs that split should confirm it inside their own property first. The likely origin of the confusion is that Google's counting documentation does describe the two features separately — a different document doing a different job from the report's dimension list.
The habit generalizes. Trade coverage of platform changes frequently runs ahead of the documentation and occasionally ahead of the product, and where a capability matters, the test is whether it appears in the operator's own reference.
Five categories of tool, and the direction each one is blind in
- Platform-owned reporting. Search Console's generative AI report and Bing's AI Performance. First-party and census-level for your own property, confined to Google and Microsoft, blind to ChatGPT, Claude, Perplexity and the Gemini app.
- Rank trackers extended to AI surfaces. Results-page scrapers that added AI Overview and partial AI Mode detection, the Semrush, Ahrefs and SE Ranking class of product. Very large keyword coverage and genuine competitor visibility, limited to surfaces that render in a results page, and sampling a logged-out approximation of what any individual sees.
- Prompt-monitoring platforms. Built to send curated prompt sets to conversational assistants and score the answers. The only way to observe ChatGPT, Claude and Perplexity at all, and the only way to measure brand mention as distinct from citation. Also a survey, driven by prompt selection at least as much as by your site.
- Log-file and CDN analysis. A census of your own server and the only view of the input side: what was fetched, how often, with what status codes. Proves nothing about whether anything was cited.
- Analytics segmentation. Measures the one quantity with direct commercial value, arriving people, and connects it to conversion. Undercounts systematically, and cannot see impressions or mentions.
No category covers the field. A program built on any one of them measures the portion of AI visibility that instrument happens to reach, and the gap is different for each.
What a vendor has to disclose before its number means anything
A visibility figure deserves trust in proportion to how much of its construction is stated: the prompt set and how it was chosen, the runs per prompt, how those runs were aggregated, the geography and account state, the collection date, and whether the surface queried was the consumer product or an API standing in for it. Those disclosures are rarely present together, and runs per prompt is the one most often absent.
The conclusion is not that the tools are worthless. It is that two vendors reporting a figure for the same brand in the same month can differ by a large multiple with neither being wrong, because they measured different prompt sets a different number of times. No conversion factor exists between them.
Movement over time carries the same defect in a sharper form. Semrush's own AI Visibility Index moved from 2,500 prompts to 126 million between successive versions. A denominator that changes by five orders of magnitude does not produce a time series, whatever the chart shows. Before comparing two dates from any tool, ask whether the prompt set and the sampling depth were identical on both; a changed prompt set invalidates the comparison however the trend line is drawn.
Frequently asked questions
Can Search Console show me AI Overview clicks?
No. Google counts those clicks — its documentation says clicking an external link in an AI Overview counts as a click — but pools them into the Web search type with no filter. The generative AI performance report launched on 3 June 2026 contains impressions only: no clicks, no click-through rate, no queries. As of 29 August 2026 there is no way to isolate them.
Does the generative AI report separate AI Overviews from AI Mode?
Google's documentation lists four dimensions for that report — pages, countries, dates and devices — and no surface dimension. Search Engine Journal reported on 4 August 2026 that the report does separate the two features. The documentation does not describe such a split, so treat any figure from the report as both surfaces combined until you can confirm otherwise in your own property.
How many times should a prompt be sampled before the result means anything?
No platform publishes a figure, and that absence matters. What is measured is the variance: 9.2% of URLs recurred across three same-day AI Mode runs on the same 10,000 keywords, and 45% of AI Overview citations change between generations. Any tool reporting a change without stating its runs per prompt is reporting variance as signal.
Is there a single number that expresses AI visibility?
No platform defines one. Vendors define their own from different inputs, over different prompt sets, at different sampling rates, which makes the resulting indices incomparable across tools and often across time within one tool. Semrush's index moved from 2,500 prompts to 126 million between releases. Ask what changed in the denominator before reading any such score as a trend.
Do GPTBot hits in my server logs mean I am visible in ChatGPT?
No. OpenAI documents GPTBot as the agent used to improve its foundation models, which is training rather than answering. OAI-SearchBot builds the index behind ChatGPT's search features, and ChatGPT-User fetches a page because a person in a conversation caused it to be fetched. Only the latter two bear on whether you can appear in an answer today.