Published
Updated
An index is a bibliography, not a body of evidence
A research index collects figures from many studies into one list. That is a useful thing to have and it is not the same thing as evidence. The index format strips each number away from the design that produced it, so a figure resting on 115 prompts sits in the same typography as one resting on 590 million searches, and a modelled forecast reads exactly like a measurement.
Two such indexes prompted this page. One collects 145 figures from a single agency's own work; the other collects 138 findings from third-party publishers. Both are honestly built, both link each figure back to its source, and neither is a study. Neither has a denominator of its own, which is why neither is graded on this site's evidence scale — an index cannot be correct or incorrect in the way a study can. What can be assessed is the corpus they point at.
One caution about counting. The first index describes itself as drawing on 21 studies, tests and case studies; its own machine-readable data resolves to sixteen source URLs, of which eight are client case studies and three are separate write-ups of a single dataset. The second describes 28 studies, but seven of the sources are the publisher's own posts, leaving 22 genuinely third-party. Neither discrepancy is deceptive. Both illustrate that a study count is itself a number with a definition behind it.
Nine companies published seventeen of the twenty-two studies
Of the 22 third-party studies in the industry index, 17 were published by companies that sell a product measuring the same quantity the study reports as alarming. Ahrefs, Semrush, Similarweb, Profound, Evertune, BrightEdge, BrightLocal, Local Falcon and SOCi all sell AI visibility tracking in some form. The five remaining come from SparkToro, iPullRank, Seer Interactive, Sterling Sky and the Pew Research Center, and of those only Pew has no commercial interest in the subject at all.
This is not an accusation and disclosed funding is not disqualifying. The vendors say who they are, and several of them publish more carefully than the people who quote them. But funding decides scope, and scope decides findings. A company selling prompt tracking measures prompt sets. A company selling backlink data measures links. A company selling local visibility scans measures the businesses already in its scanning database. In every case the question worth asking is what the study would have found had it been built to look somewhere else, and in most cases the answer is unknowable from outside.
One panel, three sources
The independence problem is sharper than the funding problem, because it is invisible in a list.
SparkToro's zero-click study draws its data from Similarweb's clickstream panel. Similarweb separately publishes two of its own studies in the same index. Presented as three rows in a bibliography, these look like three sources converging. They are one panel and three interpretations of it, and the interpretations do not agree: the same underlying clickstream supports a figure near 37% in Similarweb's own published work and 68.01% in SparkToro's, because the two draw the boundary of a click in different places.
Convergence between sources is the main reason anyone trusts a number. When the sources share a panel, that convergence is an artifact of the shared input, and an index is exactly the format in which the sharing disappears from view.
Where two vendors disagree about the same measurable fact
The clearest test of a field's evidential health is what happens when two parties measure one thing that is not in dispute.
Similarweb reports that 26% of ChatGPT responses contain an advertisement. Evertune reports roughly one in eight, which is about 12.5%. This is not a question of definition or intent; either an ad is present in a response or it is not, and it is directly countable. The two figures differ by a factor of two, neither publisher discloses the method that produced its number, and the industry quotes whichever one suits the argument being made. When a plainly countable fact cannot be pinned down within a factor of two, the more interpretive figures in the same corpus deserve considerably less confidence than they usually get.
A similar spread appears on Reddit's citation share, which runs from around 2% in one large census to 13.7% in a study of local commercial queries. That particular gap is explicable rather than contradictory, because the two measured different query universes. Which is itself the recurring lesson: no AI surface publishes a query universe, so every citation-share percentage in this field is a share of a prompt list somebody wrote.
Errors travel further than the studies they come from
Indexes propagate mistakes efficiently, because the reader who meets a figure in a list has no method in front of them to check it against.
A worked example sits in the industry index itself. It reports that 5.4% of the restaurants ChatGPT recommended are rated below 3.5 stars, and attributes this to Local Falcon's Restaurant AI Visibility Index. That study never measured ChatGPT. Its scope is Google AI Overviews and Google Maps, and the 5.4% figure describes Google's AI recommendations. The number is accurate and the engine is wrong, which is the most durable kind of error: nothing about the figure looks suspicious, so nothing prompts anyone to check it.
The same mechanism operates at larger scale on claims rather than errors. When one widely repeated finding about ChatGPT's local data sources was traced back, its origin was 50 prompts across five Spanish cities — and seven of the eight downstream write-ups had stripped every caveat within nine days of publication. The caveat is always the first thing lost, because the caveat is the part that makes the claim less useful.
How to read a research index without being misled
Four questions handle most of it, and all four are answerable in a couple of minutes from the primary source.
What is the denominator? Almost every misleading figure in this field is a correct percentage of something other than what the reader assumes. A share of citations is not a share of queries. A share of tracked businesses is not a share of businesses. A share of a keyword database is not a share of searches.
Is the comparison observed or modelled? A decline measured against a forecast is a projection dressed as a measurement, and the difference is invisible once the number is in a list.
Does the publisher sell the thing being measured? Not to dismiss the work, but to predict what it did not measure.
How many prompts? Ask before accepting any claim about what an AI engine does in general. The answer is often two orders of magnitude smaller than the confidence of the claim implies.
Where a study on this site has been read in full against its own design, it gets its own page under the studies section. This page covers the rest of the corpus, and its purpose is narrower: to record what each number is a share of, so that a figure met in a list can be returned to the design that produced it.