Published
What the Ledger counted
The AI Citation Ledger is a census of what Google's Gemini cites when it is asked to recommend a local service business. Ben Fisher published it on 6 August 2026 and last updated it on 20 August 2026. It ran a fixed matrix of the 50 largest US metropolitan statistical areas by Census-estimated population, ten service verticals, and three query phrasings, giving a target of 1,500 queries. It completed 1,487; the thirteen that failed were API errors during collection rather than sampling decisions, and 498 of the 500 metro-vertical cells carry all three phrasings.
The queries went to gemini-flash-latest through the Gemini API with Google Search grounding enabled, on 27 and 28 July 2026. The study captured the grounding chunks, the citation-to-sentence mapping, and the search strings the model generated for itself. That last item matters later. In total it classified 14,472 citations across 3,611 unique domains and extracted 4,410 unique business names.
Two details separate this from most published work in the field. The first is that it has control arms, which is rare enough to be worth stating plainly. The second is that the author audited his own classifier and published the result when it went against him. The method is documented separately and in more detail than the article itself.
The most quoted number rests on a rule the author tested and found wrong
The headline finding is that roughly 60% of Gemini's citations go to the business's own website, which is more than directories, forums and review platforms combined. That figure is a share of the 14,472 citations, not of queries and not of businesses, and Gemini emits roughly eight to ten citations per query, so a small number of citation-heavy queries carry disproportionate weight in it.
The larger issue is how a domain reaches that bucket. The taxonomy has eight categories, and business's own site is the fallback: any domain the classifier does not recognize lands there by default. Fisher checked the rule rather than assuming it. Among the fifteen highest-frequency domains defaulting into that bucket, nine were misclassified — phillymag.com and washingtonian.com were news, repairpal.com and opencare.com were lead-generation marketplaces, birdeye.com was a review-management vendor, and aaa.com and eloa.org were directories. Correcting them moved the headline from 61.1% to 58.7%.
He also states that more than 400 unique domains in the long tail have not been individually verified, and that a four-domain sample of that tail came back as genuine single businesses. So the direction of the finding survives and the precision does not. Reporting the number as about six in ten is supportable. Reporting it to a tenth of a percentage point is not, and the study's own audit is the reason.
Publishing an audit that moves your own headline down by 2.4 points is the single most credible thing in this study. It is also the thing most likely to be stripped out when the number is repeated.
The finding with a real control arm
The strongest result here is not about who gets cited. It is about whether the answer holds still.
The study asked the same queries again across five rounds five minutes apart, then compared them against the same 500 metro-vertical cells run through Google's local pack over the same schedule. Gemini returned the same top business on a byte-identical repeat 7.9% of the time, across 5,000 pairwise comparisons, with an average loose overlap of 0.177. Google's local pack returned the same top business 90.2% of the time across 4,466 comparisons, with a loose overlap of 0.884.
This is a controlled comparison against a baseline measured the same way at the same time, and it is the part of the Ledger that other work should be built on. The instability is also traced rather than merely reported: the search strings Gemini writes for itself overlap only around 5% between identical repeat calls, which places the variance in the grounding step rather than in the ranking of retrieved results. That is a mechanism finding, and mechanism findings survive longer than point estimates.
One caution the author supplies himself, and which is worth repeating because it cuts against a common reading: this is a property of generative answers, not a discovery that ordinary search is equally unstable. The control arm is what establishes that, and it is the reason the control arm needed to exist.
A study that measures a system it has proved does not hold still
Set the two halves of the Ledger next to each other and a tension appears that the study does not resolve.
One half reports citation shares to a tenth of a percentage point, with bootstrap confidence intervals from 2,000 query-level resamples. The other half demonstrates that asking Gemini the identical question twice, minutes apart, returns the same lead business 7.9% of the time. Both are correct. But a confidence interval describes sampling error around a stable quantity, and the drift result is evidence that the quantity is not stable between calls.
What the intervals honestly describe is the composition of the citation set Gemini produced on 27 and 28 July 2026 under API grounding. That is a real and useful object. It is not the same thing as the rate at which Gemini cites business websites, because there is no evidence here that such a rate exists as a fixed property to be estimated. A reader who takes 58.7% as a constant to design against has read a snapshot as a parameter.
This is a general problem with the current wave of AI search research rather than a fault peculiar to this study, and the Ledger is unusual only in containing the evidence for it.
Ratings do not predict what Gemini recommends
Every query in the matrix explicitly asked for the best or top-rated provider. The businesses Gemini named averaged 4.75 stars, against 4.84 stars for a plain Google Places category search over the same cells, with a bootstrap gap of 0.08 to 0.10 stars that does not cross zero. The AI-recommended set was slightly worse rated than what a plain search already surfaced.
The correlation between how often a business was recommended and its star rating was 0.013, which is no relationship at all. Fourteen percent of recommended businesses had fewer than fifty reviews. Ninety percent were rated 4.5 or above, which sounds like a threshold effect until the baseline is placed beside it, where 99.6% were rated 4.0 or above against the recommended set's 97%.
The correct reading is narrow and the author states it: nothing here holds market, vertical or prominence constant, so this does not prove that ratings fail to cause citation. What it does show is that among the businesses Gemini actually named, rating carried no ordering information, and that asking for highly rated providers did not produce a set rated more highly than the ordinary category listing. Anyone selling review acquisition as an AI visibility tactic has to account for that.
Reddit at 13.7%, and why another study reports 2%
Reddit alone took 13.7% of Gemini citations, with a bootstrap interval of 12.9% to 14.5%, against 10.3% for every local-service directory combined at 9.8% to 10.9%. The gap of 2.3 to 4.5 points never crosses zero. Social and community sources together reached 14.4%.
Fisher notes that a Yext study of more than 6.8 million citations put Reddit at around 2%, and he is right that this is a different measurement target rather than a contradiction. The Ledger asked one model, in one vertical family, for local recommendations, over two days. A broader citation census across query types would necessarily dilute a source concentrated in exactly the discussion-heavy commercial queries this matrix is made of.
That is the useful lesson, and it generalizes past this study. Two citation-share numbers are only comparable when the query universes behind them are comparable, and no AI surface publishes a query universe. Every percentage in this field is a share of a prompt list somebody bought or wrote.
The ChatGPT comparison measures an endpoint, not the product
The same 1,487 queries were run against ChatGPT on 30 July 2026 through a third-party scraper endpoint, then collected again on 20 August. The engines agreed remarkably little: 8.3% domain overlap, 4.2% agreement on the top business, and a loose overlap of 0.037, ranging from 0.5% in personal injury to 14.9% in auto repair.
Between the two collections, ChatGPT's social and community share moved from 41.7% to 0.0% while general directories rose from 34.6% to 40.8%. A shift that complete inside three weeks is either a product change or a change in what the endpoint returns, and the study cannot separate those two possibilities. Neither can anyone else from outside.
Both engine arms are also measured through interfaces rather than through the consumer products. Fisher says so directly about the Gemini arm, noting that API grounding is not guaranteed to retrieve or rank the way the consumer app or AI Overviews would. The comparison is between two automated interfaces, and the gap between them is a finding about those interfaces.
It shares a sampling frame with the local citations study
Steady Demand published a second and larger study in August 2026 covering AI Overviews and AI Mode, which is read separately here. The two are frequently cited together as though they corroborate each other. They are independent collections — different engines, different tooling, different dates, different volumes — but they are built on the same 50-metro by 10-vertical scaffold, and the AI Overviews study describes the Ledger as its companion.
Sampling the same 500 metro-vertical cells twice is not replication. Any bias in the frame, and a top-50-metro US home-services frame has several, is present in both results in the same direction. Where the two studies agree, that agreement is weaker evidence than it appears. Where they disagree, as they do sharply on how much of the citation set goes to Google's own properties, the disagreement is informative precisely because the frame is held constant and the engine is not.
Who published it, and what the funding decided to measure
Steady Demand is a local SEO agency that sells AI optimization for local businesses, Google Business Profile management and reinstatement, and Local Services Ads management. It also runs a free AI optimization audit tool. No external funding is disclosed, and the study carries no explicit conflict statement.
The commercial interest shows in scope rather than in conclusions. The matrix is home services and local professional verticals, which is the client base; the surfaces are the ones those clients ask about; the metrics are the ones an agency would report. What the study is not built to answer is whether any of this is actionable, and the methodology page says as much in a line that works against its own author's sales pitch: a system that does not hold still should not be sold with a flat promise that doing X produces a ranking in AI.
One loose end is worth recording. The public dashboard and downloadable data accompanying the Ledger describe coverage of five of fifty metros and 149 queries, against an article reporting 1,487 queries across all fifty. The published extract therefore appears to cover roughly a tenth of the dataset the findings rest on, and the raw-data page states no row count, schema or license. The figures cannot be independently recomputed from what has been released, which is a limitation on verification rather than a criticism of the analysis.
Frequently asked questions
What did the AI Citation Ledger measure?
It measured which sources Google's Gemini cites when asked to recommend local service businesses. It ran 1,487 queries built from the 50 largest US metros, ten service verticals and three phrasings through the Gemini API with Google Search grounding on 27 and 28 July 2026, and classified 14,472 resulting citations into eight source categories.
Is it true that 60% of Gemini's citations go to business websites?
Approximately, and the study's own audit is the reason to say approximately. Unrecognized domains default into the business-website category. Checking the fifteen most frequent defaults found nine misclassified and moved the figure from 61.1% to 58.7%, and more than 400 long-tail domains remain individually unverified. About six in ten is supportable; a precise figure is not.
What is the strongest finding in the study?
The stability comparison. Gemini returned the same top business on an identical repeated query 7.9% of the time, against 90.2% for Google's local pack measured over the same cells on the same schedule. It has a genuine control arm, and the overlap of the model's self-generated search strings locates the instability in the grounding step rather than in ranking.
Does the study show that reviews and ratings do not matter for AI visibility?
No, and the author says so. Recommended businesses averaged 4.75 stars against a 4.84-star plain-search baseline, and the correlation between recommendation frequency and rating was 0.013. But nothing was held constant, so this cannot isolate the effect of rating. It shows that among the businesses Gemini named, rating carried no ordering information.
Can this study be combined with the Steady Demand AI Overviews study as corroboration?
Not as independent corroboration. The two are separate collections but are built on the same 50-metro by 10-vertical sampling frame, and the AI Overviews study calls this one its companion. Shared-frame results repeat any bias in the frame rather than testing it.
Why does this study report Reddit at 13.7% when others report about 2%?
Because the query universes differ. This matrix is entirely local commercial recommendation queries, where discussion sources concentrate. A broad citation census across all query types dilutes that share. No AI surface publishes a query universe, so every citation-share percentage is a share of the prompt list its authors chose.