A technical reference on AI search visibility. This site sells nothing, takes no engagements and endorses no products. Consulting enquiries are handled separately at hartzer.com.

Hartzer.it.com logoHartzer.it.comAI search visibility reference
Abstract tapered column illustration representing How to evaluate an AI visibility tracking tool
Measuring AI Search Visibility

How to evaluate an AI visibility tracking tool

No platform defines AI visibility, so every score is a vendor's own construction. The question is not which tool is best but what its number was made from.

UnsupportedNo platform defines AI visibility, and no vendor was found publishing the prompt set or sampling depth needed to reproduce its score.

Start from the question, not from the category

Tools in this category are bought as though they were interchangeable, and they are not. They answer different questions, and several of the questions buyers care about most have no tool behind them at all. Before comparing products, write down which of these you actually need answered.

  • Are links to my pages being shown inside Google's generative answers? Google's own reporting answers this for your property, as a count rather than a sample.
  • Is my brand named when someone asks an assistant about my category? Only a prompt-monitoring product can observe this, and only by sampling.
  • Is my share of the citation pool moving? One platform publishes a share metric with a real denominator. Every other share figure is a share of somebody's prompt list.
  • Did anyone arrive, and did they do anything? Your own analytics, and only for the surfaces that send an identifiable referrer.
  • Were my pages fetched at all? Server and CDN logs, which are the only census of the input side.

No single product covers that list, and each category of tool is blind in a specific and predictable direction. Platform reporting misses everything outside Google and Microsoft. Rank trackers extended to AI surfaces miss everything that does not render in a results page. Prompt monitoring is a survey. Log analysis measures input rather than output. Analytics measures the small surviving fraction of output. A program built on one of them will be confidently wrong in the direction that instrument cannot see.

The survey at the center of the category

Almost every product marketed as an AI visibility platform works the same way: it sends a curated list of prompts to each assistant on a schedule, parses the answers, and records whether the brand was named and whether the domain was cited. That is a survey, and it inherits every property of one: the result depends on which prompts were chosen, how many times each was run, from where, on which surfaces, and in what account state.

The prompt set is the single largest determinant of the output, and it is almost always proprietary. This is worse than it sounds, because there is no way to build a representative one. No AI operator publishes prompt volume data, so there is no equivalent of search volume to weight against, and prompt sets are assembled from keyword research, from client input, or by asking a language model to generate plausible questions — all reasonable, none of them a sample frame. A number described as your visibility is your appearance rate in a list of questions a vendor wrote, and a buyer who does not see the list cannot audit the number even in principle.

That is not an accusation of bad faith. It is a statement about what the instrument is. Treat these products as survey research, ask the questions you would ask any pollster, and they become genuinely useful. Treat them as a meter reading and they will mislead you in ways that look like performance.

Why two honest tools disagree about you

Buyers routinely find that two products report different visibility for the same brand in the same month and conclude one of them is broken. Usually neither is. Four independent choices separate them, and any one alone can move a score by a large factor.

Surface coverage. The surfaces overlap far less than their similar answers suggest. Ahrefs measured 13.7% citation overlap between AI Overviews and AI Mode across 540,000 query pairs in December 2025, despite the answers being highly similar in meaning; SE Ranking put the URL-level overlap between the two at 10.7%. Choosing which surfaces to include therefore changes the answer before any other decision is made.

Sampling depth. Rarely disclosed, and decisive, for the reason set out below.

What counts as an appearance. Some products count brand mentions in the answer prose, some count domain citations, and some blend the two at an undisclosed weight.

What the tool actually queried. Some products call an API rather than the consumer product. Those are different systems with different retrieval behavior, and a result from one is not evidence about the other.

The consequence for procurement is blunt. Two vendors' visibility scores for the same brand are not comparable, there is no conversion factor between them, and a competitive comparison assembled from two different tools is not a comparison. If a pitch shows your score against a rival's, ask whether both numbers came from the same instrument, on the same prompt set, in the same window.

Mentions and citations are two measurements

A brand mention in the answer prose and a citation of your domain as a source are produced by different mechanisms and come apart in both directions, which is why the better-built products keep them separate and the weaker ones average them into a single figure.

Seer Interactive named the first failure in March 2026: the ghost citation, where an answer uses your page as a source and links to it without ever speaking your brand name. Your content did the work and your brand got none of the credit. The reverse happens too — an answer recommends you by name and links somewhere else entirely, which produces no citation, no referral, and quite possibly a customer. Semrush's June 2026 index deliberately reports the two as distinct metrics, defining mentions as how often a company appears in an answer and citations as which domains and pages the platforms use as evidence, and reports the two sets diverging sharply on Gemini.

A blended visibility score hides which of the two moved, and they call for opposite responses. A mention problem is an entity and reputation problem that lives largely off your own site. A citation problem is a retrieval and content problem that lives on it. If a dashboard cannot decompose its own headline number into those two, it cannot tell you what to do next, which is the only reason to buy it.

Exhaust the census before paying for a sample

Two platforms report on their own output, and both are free to the site owner. They are narrower than the commercial products and they are not samples, which for some questions makes them strictly better evidence.

Google Search Console's generative AI performance report, live since 3 June 2026, gives impressions of links to your site shown in a generative AI feature on Google Search, by page, country, date and device. No clicks, no click-through rate, no queries, and no split between AI Overviews and AI Mode.

Bing Webmaster Tools went further and earlier. Its AI Performance report, in public preview since 10 February 2026, reports citation counts across Microsoft Copilot, AI-generated summaries in Bing and select partner integrations, and it exposes the grounding queries — the retrieval queries the system actually issued, as opposed to what the user typed. That is the sub-query layer Google exposes nowhere, and for anyone trying to work out why a page was cited it is the most informative artifact any platform publishes. Since 16 June 2026 Microsoft has also published Citation Share, defined as the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query (Bing Blogs).

Microsoft is also candid about the limits, and the candor is worth more than the metric. Citation data "does not indicate ranking, authority, or the role of any page within an individual answer"; Citation Share "does not expose competitor domains" and is not a "competitive scoreboard"; and there is no traffic metric anywhere in the report. Bing's reach is smaller than Google's. Its instrumentation is better.

Where the method breaks

Five failure modes account for most of the bad numbers in this category, and a tool's handling of them tells you more than any feature list.

  • Single sampling. SE Ranking parsed the same 10,000 keywords three times in one day in AI Mode and found 9.2% of cited URLs present in all three runs (SE Ranking, published 29 August 2025). Ahrefs found 45% of AI Overview citations changing between generations. A product that reports a week-on-week change without disclosing runs per prompt is reporting variance as signal.
  • A prompt set that changed under the chart. Semrush's own visibility index moved from 2,500 prompts to 126 million between releases. A series whose denominator changed by five orders of magnitude is not a series, however smooth the line looks.
  • Personalization. Since 27 May 2026, Preferred Sources selections are reflected in AI Overviews and AI Mode, so part of the citation set is configured by the reader. Semrush's own methodology note concedes that AI Mode may tailor responses to user context and that citations may vary from person to person.
  • One turn standing in for a conversation. John Mueller confirmed on 6 August 2026 that a follow-up question inside AI Mode is a new query, with its own impressions, position and clicks. A tool that issues one prompt is not measuring what a user who asks four is doing.
  • Invented ordering. Any score weighted by where a citation sits in the answer contradicts both platforms that have commented on it.

The questions to ask before believing a number

These are the disclosures that separate a measurement from a guess. Ask them in writing, and treat a refusal as information rather than as an inconvenience. A vendor that answers all of them has earned a place in a report; one that answers none of them is selling a chart.

  • What is the prompt set, how large is it, and who chose the prompts?
  • How many times is each prompt run inside a reporting period, and how are the runs aggregated into the number I am shown?
  • Did the prompt set change between the two dates this chart compares?
  • Which surfaces are covered, and for each, is the tool querying the consumer product or an API?
  • From which country, on which device profile, and in what account state?
  • Is the headline number mentions, citations, or a blend — and if a blend, at what weight?
  • What is the denominator: all prompts, or only prompts in which some brand appeared?
  • What is the confidence interval, or the observed run-to-run variance for my own brand?

The single most informative answer is runs per prompt. It is almost never on the dashboard, and everything else in the methodology is downstream of it.

What no tool in this category can do

Some questions have no instrument behind them, and a program that assumes otherwise will spend money looking for data that does not exist.

No tool can show you clicks from Google's generative surfaces. Google counts them and pools them into the Web search type, where they cannot be filtered, and no third party can see what Google will not expose. No tool can show you Google's fan-out sub-queries, because Google publishes them nowhere. No tool can attribute what is plausibly the most common outcome of an AI answer, the brand named without a link, because that visit has no referrer to lose. And no tool can show that a change on your site caused a change in citations, since none of them run a holdout; the one systematic review of the field, Martinez's survey of 45 studies published 15 July 2026, found no reviewed technique with a stable, longitudinal, cross-platform causal effect on discoverability.

And no tool can produce a number comparable to another tool's. There is no standard definition of AI visibility from any platform, standards body or industry group, and as of 29 August 2026 no vendor was found publishing a formula, prompt set or sampling depth in a form that would let anyone reproduce its figure. That is recorded here as an absence rather than as an accusation, but it has a hard consequence: no commercial AI visibility score on the market can currently be independently checked, and a number that cannot be checked should be labeled as such wherever it is reported.

Frequently asked questions

Is there a standard definition of AI visibility?

No. As of 29 August 2026 no platform, standards body or industry group publishes one, and no two vendors were found publishing the same formula. The phrase carries an air of rigor borrowed from advertising metrics that do have agreed definitions, and the inheritance is unearned. Ask each vendor for its own definition in writing and expect them to differ.

Can I compare my score in one tool with a competitor's score in another?

No, and the comparison is not merely imprecise, it is undefined. The two numbers use different prompt sets, different surfaces, different sampling depths, different numerators and different denominators. With as little as 13.7% citation overlap between AI Overviews and AI Mode, surface selection alone can produce a large divergence before any other choice is made.

Do I need a paid tool at all?

It depends entirely on which surfaces matter to you. Google and Microsoft each report on their own output for free, and for Google-heavy and Bing-heavy questions that reporting is a census rather than a sample. Nothing free covers ChatGPT, Claude, Perplexity or the Gemini app, because their operators publish no reporting for publishers at all. Prompt monitoring is the only way to observe those, and its limits are the limits of survey research.

Is a tool that queries an API measuring the same thing as the consumer product?

It is measuring a related system, not the same one. The consumer products apply their own retrieval, personalization and interface behavior, and no operator publishes how closely an API response corresponds to what a user sees. Ask which one the tool queried and record the answer beside the number.

What single disclosure matters most?

Runs per prompt. Given 9.2% three-way URL reproducibility in AI Mode and 45% of AI Overview citations changing between generations, sampling depth determines whether a reported movement is a finding or noise. Everything else in a methodology is secondary to it, and it is the disclosure least often volunteered.

How should a tool report change over time?

As a rate over a frozen prompt set, with the run count and the collection window printed beside it, and with a note whenever the prompt set changed. Semrush's index moving from 2,500 to 126 million prompts between releases (Semrush) is the clean illustration of why: the two releases measure different things and cannot be plotted as one line.

Top