Published
A term that arrived from SEO with its evidence left behind
Topical authority is the idea that covering a subject thoroughly, rather than chasing individual queries inside it, makes a site the one a search system trusts on that subject. It entered SEO vocabulary in the mid-2010s, alongside the shift from keyword matching to entity-based retrieval, and it has been useful advice for a long time.
It was never a documented Google ranking factor. Google has described relevance, quality and the entity relationships in its knowledge systems; it has never published a topical authority score, and practitioners who use the term are describing an outcome they infer from behavior rather than a mechanism anyone has confirmed.
The concept has now been carried across to AI search largely intact, and the carrying is the problem. In organic search the claim was tethered to something a practitioner could check: rankings for a cluster of queries, measured against coverage of that cluster. In AI search the outcome being claimed is being named in an answer, which is produced by a different part of a different system, and the tether did not come along.
What follows is what is actually known, which is less than the term implies and more than nothing.
What the operators document, which is relevance and nothing named
No AI operator has published a topical authority signal. Google, OpenAI, Microsoft and Perplexity each describe eligibility conditions for a page to be usable as a source — indexation, accessibility to the relevant crawler, relevance to the query, and quality language of varying specificity. None of them describes a measure of how thoroughly a domain covers a subject, and none reports one back to publishers.
Relevance, which every operator does name, is not the same claim. Relevance is evaluated for a document against a query. Topical authority is a property asserted about a domain across a subject. A system can score the first without ever computing the second, and nothing in the public documentation indicates that the second exists as a distinct input.
Bing Webmaster Tools reports citations and citation share for grounded queries. Search Console reports impressions that include AI surfaces. Neither reports anything resembling subject coverage, so a publisher cannot observe this variable in first-party data at all. Every figure in circulation comes from a third party sampling prompts.
The gap matters because of what fills it. In the absence of a reported signal, the industry substitutes a proxy: how many pages a site has published in a subject, or how tightly those pages cluster in a vector space, or a vendor's composite of the two. Each of those is a description of the publisher's own output. None is a description of anything the system does with it, and treating a proxy as the signal is how a measurable input gets mistaken for a documented one.
Perplexity is the useful test case here. It makes citation the visible product, shows its sources on every answer, and still does not explain selection. If any surface were going to expose a subject-level source preference, it would be the one whose interface is built out of sources. It does not.
The one panel that has measured it
Semrush, working with Growth Memo, ran the only large measurement of the question that exists as of September 2026. Two studies on the same panel of 1,094 US category prompt sets, tracked in ChatGPT from January through June 2026.
The first found that domain-level metrics — branded search volume, Authority Score, organic traffic — separated category leaders from runners-up at close to chance, which removed the obvious alternative explanation without supplying one.
The second supplied a candidate. Proximity between a new category and the categories a brand already answered well tracked the outcome that the domain metrics had not. Near a brand's established subjects, 74% of its appearances came with a citation and 44% with a named mention. At the far end of the range, 50% and 25%. The share of citations that also carried a brand name ran 46% in the most related categories against 18% in the least related.
That is a real association, replicated across two analyses of one panel, and it is the strongest evidence the concept has ever had on any surface. It is also a single vendor's prompt list, measured without controls, on a system that does not answer the same question twice the same way.
The two analyses are covered individually: the topic authority study for the ownership distribution and the near-chance result on domain metrics, and the topical focus study for the distance curve and the depth threshold. Both are read there against their own design rather than summarized.
The threshold finding, and why it is the fragile one
The most actionable claim to come out of that work is a threshold. Brands appearing in only one of a category's five prompts were associated with a drop in mention share, and the penalty did not clear until three of five.
Taken at face value this inverts a piece of long-standing practice, because publishing one piece into a new subject to test it has always been cheap and harmless. A threshold of this shape would mean partial coverage costs something.
It is also the finding most exposed to the direction of causation. A brand appearing in one prompt out of five is plausibly a brand few buyers could name unprompted. Its mention share would be low whatever it published. On a cross-sectional design there is no way to distinguish shallow coverage causing weak recognition from shallow coverage being what weak recognition looks like in a category.
The expansion analysis behind it also rests on 1,458 mapped brand entities, against the 50,000-plus brands in the first study on the same panel. The threshold is worth planning against. It is not worth quoting as a rule.
The mechanism that is real, and the promise attached to it
There is a coherent reason depth would matter, and it is not a topical authority score.
Being cited is a retrieval outcome: a document matched a query and cleared the source conditions. Being named is a generation outcome: the model produced a brand because its representation associates that entity with that subject. The first responds to what a publisher publishes. The second responds to what the corpus says about that publisher, which is mostly written by other people.
Deep coverage plausibly feeds both — more documents to retrieve, and a stronger subject association in whatever the model absorbed. That is a reasonable inference from documented retrieval behavior and from how entity representations are built. It is not a measurement, and it does not license the step most vendors take next, which is to treat published coverage as the input that moves the named-mention outcome.
The distinction has a practical edge. If mentions respond mainly to the corpus rather than to a publisher's own pages, then a program that funds only content production is funding the signal that travels between categories and not the one that gates them.
What gets sold on top of this
The commercial version of topical authority in AI search is content volume with a taxonomy attached. Cover every question in a cluster, publish across as many adjacent clusters as production allows, and visibility follows.
The measured panel argues against the second half of that. Breadth without depth was the configuration associated with the worst outcome in the data, and citation counts — the metric a volume program most easily moves — were the signal that spread across categories regardless of standing. A program can raise its citation chart and leave the mention chart flat, and on this evidence that is the expected result of publishing wide.
The other common sale is a topical authority score. Several vendors compute one. None of them is measuring a platform signal, because no platform exposes one; each is a construction over that vendor's own crawl and prompt panel, and two vendors scoring the same domain can disagree without either being wrong about its own inputs. A score built on a private denominator is a description of that vendor's sample.
A third pattern is worth naming because it is the most persuasive of the three. Case studies pair a coverage build-out with a rise in AI visibility over the same months and present the sequence as the mechanism. The problem is not that the numbers are invented; usually they are not. It is that a company investing heavily in subject coverage is also, in the same quarter, doing public relations, shipping product, running campaigns and accumulating third-party coverage. All of those feed the corpus a model reads. A single-site before-and-after cannot separate them, and none of the published case studies attempts to.
What would settle it
The experiment is straightforward and nobody has run it publicly.
Take a set of brands with comparable recognition inside a category, established by a measure independent of the platform — branded search volume, survey recall, or existing mention share as a baseline. Assign coverage depth at random: one prompt's worth of content, three prompts' worth, five. Hold the writing quality, the publication cadence and the off-site activity constant. Sample each prompt enough times to get past non-determinism, and measure mention share before and after.
Anything less than that leaves the direction open, which is where the question stands today. Two observational studies of one panel, no controls, no replication by a second party, and no platform confirming that subject coverage is an input at all.
The honest working position is that depth in a subject is a reasonable thing to invest in for reasons that predate AI search, that it accompanies AI visibility in the only data that exists, and that anyone telling you it causes AI visibility is telling you something nobody has tested.
Frequently asked questions
Is topical authority a ranking factor in AI search?
No operator has named one. Google, OpenAI, Microsoft and Perplexity all describe relevance as an eligibility condition for a document against a query. None describes a measure of how thoroughly a domain covers a subject, and none reports one to publishers.
What is the evidence that deep coverage helps?
Two 2026 analyses by Semrush and Growth Memo on one panel of 1,094 ChatGPT category prompt sets. They found that proximity to a brand's existing subjects tracked whether it was named in a new category, and that presence in fewer than three of five prompts was associated with a lower mention share. Both are cross-sectional and neither tested direction.
Should I publish into fewer categories and go deeper?
That is the reading the measured panel supports better than the alternative, since breadth without depth was the configuration with the weakest outcome in that data. It is a planning judgment on observational evidence, not a demonstrated result, and it is worth holding as such.
Is a vendor topical authority score measuring anything real?
It is measuring that vendor's crawl and prompt panel. No platform exposes a topical authority signal, so any score is a construction over private inputs. Two vendors can rate the same domain differently and both be internally consistent.
Does topical authority in AI search mean the same thing as in organic SEO?
The term is the same and the outcome is not. In organic search it describes rankings across a cluster of queries, which a practitioner can observe. In AI search it is usually applied to being named in an answer, which is produced by generation rather than by document ranking, and which no first-party reporting surface exposes.
What would count as proof?
A randomized comparison: brands matched on prior recognition, coverage depth assigned rather than observed, writing quality and off-site activity held constant, and enough sampling per prompt to clear non-determinism. Nothing of that design has been published.