Published
What the study counted
Semrush tracked 1,094 US categories in ChatGPT every month from January through June 2026, drawing on more than 220,000 domains and over 50,000 brands recorded in its AI Visibility Toolkit. A category here is not a keyword. It is a cluster of five prompts written to follow the sequence a buyer moves through: what is this, how do the options compare, what else is out there, does it fit my situation, and should I buy it. Payroll software is a category. Small business banking is a category.
For each of those prompts the researchers logged two separate things: which brands ChatGPT named in the prose of its answer, and which sources it linked underneath. The study then built its findings on the first of those and largely set the second aside.
That choice carries the whole paper, so it is worth stating plainly. Semrush scored brands on mentions in the answer text rather than on citations, justifying it with earlier research finding that 74% of users chose the top-mentioned brand as their final pick. That figure comes from outside this study. It is a borrowed statistic doing structural work, and a reader who rejects it is left with a paper measuring something whose commercial meaning has not been established here.
Three classifications came out of the tracking. A clear owner held the highest mention share, appeared in at least four of a category's five prompts, and led the runner-up by five percentage points or more. An emerging leader reached three prompts without meeting that bar. A category where no brand reached three prompts was unsettled. The original is published as the Semrush ChatGPT topic authority study.
Three buckets, and the inversion inside them
Across all 1,094 categories, 15.2% had a clear owner, 31.2% had an emerging leader, and 53.7% were unsettled. That last number is the one that travels, usually rendered as proof that AI search is wide open.
The more interesting result is what happened when the categories were split in half by AI search volume. Among the top half by volume, only 11.3% had a clear owner. In the bottom half, 19% did. And the top half carries 98% of all the AI search volume in the panel.
So the commercially valuable categories are less settled than the obscure ones, which reads at first as a straightforward opportunity finding and should not.
A high-volume category attracts more competitors. More competitors split mention share more ways, and a field of eight brands makes a five-percentage-point lead arithmetically harder to hold than a field of two. Part of that 11.3% is crowding rather than vacancy. The study does not separate the two, and nothing in its design lets it: brand count per category is not reported alongside the ownership rate, so a reader cannot tell how much of the difference is competitive density.
The practical reading survives either interpretation. In the categories where the money is, this panel found no brand consistently holding the answer. Why that is true matters for what you do about it.
The domain metrics landed on chance
Having identified owners, the study asked what separated them from their runners-up, and tested three familiar site-level measurements against paired head-to-head comparisons.
Owners had higher branded search volume in 55.7% of pairs. A higher Authority Score in 52.5%. Higher organic traffic in 48.4%, which is a shade worse than a coin toss.
None of those separates anything. Kevin Indig, who worked on the study, was careful about the conclusion people would reach:
Treat that as a hypothesis, not a finding. What the data actually shows is: traditional SEO metrics aren't enough to explain who owns a topic.
That caution is the right posture and it is routinely dropped in secondhand coverage, which converts a near-random result into the claim that SEO no longer matters in AI search. The study does not support that claim. It reports that three specific domain-wide numbers do not predict a category-level outcome, and offers no evidence about any other input.
Why a near-random result was the expected outcome
There is a structural reason those three comparisons landed where they did, and it is not that search fundamentals stopped working.
Branded search volume, Authority Score and total organic traffic are properties of an entire domain. Category ownership as this study defines it is a property of five prompts. Testing one against the other compares a whole-site aggregate to a single-cluster outcome, and an aggregate is precisely the thing that cannot resolve variation inside itself.
Any technical audit of a large site demonstrates the same point without needing a panel. A domain can carry excellent sitewide numbers and still hold four thin pages against the one subject its revenue depends on. That gap is invisible in every metric the study tested, by construction, because averaging is what those metrics do.
The correct conclusion from the 48.4% figure is therefore narrower and more useful than the one in circulation. It is not that authority is irrelevant to AI search. It is that domain-level measurements are the wrong unit of analysis for a category-level question, and any study using them should expect results near chance. A comparison at matching granularity, such as coverage depth within the category itself, is a different test, and this study did not run it. The follow-up research did.
The persistence finding, and the only threshold here worth using
The most operationally useful part of the study is the part with the least commentary attached to it.
Brands classified as clear owners held first place in 90.4% of month-over-month comparisons. Among emerging leaders and unsettled categories, leadership changed hands in 1,950 of 5,470 comparisons.
The margin data explains the difference. Categories where the leader flipped carried a median lead of 1.3 percentage points. Categories where the leader held carried a median lead of 2.9.
That gives a reader something no vendor dashboard supplies: a rough noise floor. A one-point lead in mention share is not a position, it is measurement variance, and treating it as progress is how a monthly report manufactures a trend out of a non-deterministic system. Somewhere near three points the lead starts to persist. At the five-point owner threshold it persisted nine times out of ten in this panel.
Two caveats belong with that. The thresholds are specific to this panel and this prompt design, so they are a reference point rather than a constant. And persistence in a six-month window says nothing about whether a position is defensible against a competitor who decides to contest it, because the study observed categories rather than interventions.
The denominator is a list Semrush chose
Every percentage in this study is a share of 1,094 categories that Semrush selected and five prompts per category that Semrush wrote. No AI surface publishes a query universe, so there was no alternative available. That does not make the figures wrong; it makes them conditional on a list, in the same way every vendor share-of-voice number is conditional on a list.
The consequence is that the headline distribution is not a property of AI search. It is a property of this panel. A panel weighted toward software and financial services categories will report a different ownership rate than one weighted toward consumer goods, and the study does not publish the category composition needed to tell how much of the 53.7% is a sampling artifact.
Five prompts per category is the second design limit, and it interacts badly with a system that does not answer the same question the same way twice. Under a four-of-five threshold for ownership, a single anomalous run demotes a brand from owner to emerging leader. Non-determinism therefore pushes categories down the classification scale rather than scattering them in both directions, which means the unsettled share is better read as a ceiling than as a floor.
The panel is also ChatGPT alone, in the United States, over six months. Nothing in it speaks to Gemini, Perplexity, Copilot or Google's AI Overviews.
What this study supports, and what it does not
It supports a descriptive claim: within one vendor's panel of 1,094 category prompt sets, most categories had no brand ChatGPT named consistently, high-volume categories were less settled than low-volume ones, and brands that did establish a wide lead usually kept it across the following month.
It supports a negative claim about method: three specific domain-level metrics failed to separate category owners from runners-up, so a visibility program that reports those metrics is not reporting on category ownership.
It does not support any causal claim. There was no intervention and no control group, so nothing here establishes that acquiring branded search volume, authority or traffic would change a brand's mention share. It does not support cross-platform generalization. It does not establish that mention share drives revenue; that link is imported from separate research on user choice, not measured here.
Treated as a first map of an unmeasured territory, it is the most useful thing published on category-level AI visibility to date. Treated as a ranking-factor list, which is how it is already being summarized in places, it is being asked to carry weight its design cannot bear.
Frequently asked questions
What did the Semrush topic authority study actually measure?
How often ChatGPT named specific brands in the text of its answers across 1,094 US category prompt sets, tracked monthly from January to June 2026. Each category was five prompts covering a buyer's sequence of questions. The study counted brand mentions in the answer prose, not citations in the source list.
Does 53.7% unsettled mean most AI search categories are open?
It means 53.7% of the categories in this panel had no brand appearing in at least three of five prompts. Because a non-deterministic system will occasionally drop a brand from one prompt, and the classification thresholds are unforgiving in one direction only, that figure is better read as an upper bound on fragmentation than a lower one.
Did the study show that SEO metrics no longer matter?
No. It showed that three domain-level metrics predicted category-level ownership at close to chance. Those are different units of measurement, and a whole-site average cannot resolve variation between subjects inside that site. Kevin Indig described the broader interpretation as a hypothesis rather than a finding.
Is the 90.4% retention figure a reason to move quickly?
It is evidence that wide leads persisted month to month within this panel over six months. It is not evidence that a position is defensible against a competitor contesting it, because the study observed categories rather than interventions. The margin data underneath it is more useful: leads flipped at a median of 1.3 points and held at 2.9.
Can this study's percentages be compared with another vendor's?
Not directly. The denominator is Semrush's own list of categories and prompts, and no AI surface publishes a query universe that would allow two vendors to divide by the same thing. Two studies can both be internally correct and disagree by a multiple.
Who funded the research?
Semrush, which sells AI visibility tracking, produced it with Kevin Indig of Growth Memo. The funding is disclosed rather than hidden, and it shapes what got measured: a company selling prompt-set tracking measured prompt sets. That is worth holding in mind when reading which questions the study asked and which it did not.