Published
Expansion is a different question from coverage
Whether depth in a subject helps at all is one question, and it has its own page. Category expansion is the next one: given that a brand is established in some subject, how much of that standing carries into a neighboring subject, and how near does the neighbor have to be?
The distinction is not academic. A content plan is not usually a choice between publishing and not publishing. It is a choice of which subject comes next out of a list of candidates that all look plausible, and the usual tiebreaker is search volume. Volume tells you how big a category is. It says nothing about whether you can win it, and until recently nothing did.
The measurement that exists treats adjacency as the variable. A category is close to a brand when the questions asked inside it resemble the questions the brand already answers well; distant when they do not. That framing makes the decision testable in a way that a taxonomy of industries never was, because it is computed from the prompts themselves rather than from a category tree somebody drew.
One framing note before the evidence. Everything below describes what accompanies successful expansion in observational data. None of it describes an intervention anyone has run, so the useful posture is that of a practitioner reading a well-built map of terrain nobody has yet walked with instruments.
The two signals decay at different rates
The central finding is an asymmetry. As a category gets further from a brand's established subjects, its chance of being cited as a source falls moderately. Its chance of being named in the answer falls much faster, and the chance of both happening together falls fastest of all — from roughly a third of appearances at the near end of the range to under one in ten at the far end.
A brand can therefore be present in a distant category, see its citation counts rise, and remain effectively invisible on the measure that corresponds to a buyer hearing its name. Any dashboard reporting one aggregated visibility figure will show that as progress.
This is the same decoupling that shows up in citation and mention data generally, and category distance appears to be a large part of what drives it. Pooled across all distances the relationship between being the most-cited domain and the most-named brand is near zero. Split by distance it has clear structure. A near-zero average across a structured distribution is a familiar statistical shape and a misleading one to act on.
There is a second reason the aggregated figure misleads, and it is arithmetic rather than statistical. Citations are far more numerous than named mentions in every panel that has counted both — by close to four to one in the largest of them. A composite visibility score that sums the two, or weights them by frequency, is dominated by the signal that travels and barely registers the one that gates. The metric moves most where the brand has least standing.
Distance is measurable, and the measure is ordinary
The distance construct is not exotic. Take the set of questions that define a category, embed them with any current text embedding model, do the same for the categories a brand already answers well, and compute cosine similarity between the two sets. Rank the candidates by that score.
An in-house team can reproduce a workable version of this in an afternoon, which is worth saying plainly because the alternative on offer is usually a subscription. The inputs are a list of candidate categories, five or so representative buyer questions each, and an embedding endpoint. Nothing about it requires the vendor's panel.
Two cautions come with it. Similarity computed on question text measures vocabulary overlap, and vocabulary overlap is a proxy for market adjacency rather than the thing itself. Two categories can share language and no customer. The filter that catches this is ordinary commercial judgment: genuine adjacency usually means a shared customer, a shared problem, a product that plausibly covers both, and a comparable buying process. Where a high similarity score fails those, the score is picking up words.
The second caution is that the reference figures below were produced on one vendor's prompt design. Your own similarity scores will not be on the same scale, so they rank candidates rather than predict outcomes.
The conversion ratio is the diagnostic
The most useful instrument to come out of this work is a ratio rather than a count: of the citations a brand earns in a category, what share also carry a named mention of it.
It is worth more than either raw number because it is insensitive to publishing volume. A site can lift its citation count in a subject indefinitely by publishing more, and the ratio will not move unless recognition moves with it. That makes it a check on the failure mode this page is about — activity in a category that produces sources without producing standing.
The published reference points, from the one panel that has measured them, are about 46% in the categories most related to a brand's core and about 18% in the least related. A new category where the observed ratio sits near the low end is evidence that the content arrived somewhere the audience-facing signal has not followed.
No platform reports this. Bing Webmaster Tools reports citations without mentions; Search Console reports impressions and neither of the two. Assembling the ratio requires a tool that separates the counts by category, and the number it produces inherits that tool's prompt list along with every other limit that carries.
Sector sets a ceiling on what expansion can do
The relationship is not uniform across industries. In the measured panel, financial services and real estate converted broader coverage into citation leverage. Legal and healthcare did not, and in those sectors the mention trend stayed negative even for brands present across an entire category's question set.
That is a ceiling rather than a slope. It says that in some subjects, coverage is the entry requirement and not the lever, and that a plan built entirely on publishing will hit a limit that more publishing does not move.
What produces the ceiling has not been tested. The pattern resembles the raised bar that ranking systems have applied to health and legal subjects for a decade, and the obvious candidates are credentialing, licensure and independent validation. The studies that found the ceiling did not measure author credentials, primary research or product specifics, and said as much. Anyone in those sectors should treat the ceiling as observed and the explanation as an open question.
The practical consequence is a different budget shape in those sectors. If publishing cannot move the ceiling, the spend that might is the spend on the things the studies did not measure: named authors with verifiable qualifications, original data nobody else has, and coverage by publications and reviewers with their own standing. That is a communications and credentialing investment rather than a content one, and it sits in a different part of most organizations.
What the pattern implies for sequencing
If the measured relationships hold, expansion has a shape rather than a budget.
It starts from anchors: the categories where a brand already earns both citations and named mentions across most of the question set. Anchors are the launch points, and most organizations have fewer of them than their content inventory suggests.
It proceeds outward one step at a time, because the variable that tracked the outcome is proximity to an existing anchor, and a category only becomes an anchor once it is answered well. Two steps out is not reachable by skipping one.
It enters at depth rather than as a test. The threshold reported in the panel puts the floor at three of a category's five buyer questions, and below it partial presence was associated with a worse outcome than staying out. That converts category entry into a minimum commitment, which is a scheduling constraint more than a spending one: the same budget buys three categories done properly or twelve done shallowly, and the two are not equivalent on this evidence.
And it is measured on the ratio rather than on the counts, because the counts move for reasons that have nothing to do with whether the expansion worked.
Those four moves are the substance of planning topic coverage as a sequence rather than as a list, and the step-by-step version, including how to sample against non-determinism and what a monthly report should carry, is set out separately.
What is measured, what is inferred, and what is neither
Measured, within one vendor's panel of ChatGPT prompts over six months: citations decay modestly with topical distance, named mentions decay sharply, the conversion ratio between them ranges roughly from 18% to 46% across that distance, presence below three of five prompts is associated with a lower mention share, and two regulated sectors do not follow the general pattern.
Inferred, from documented retrieval behavior rather than from the study: that citations travel because retrieval scores documents while naming reflects an entity association built from the whole corpus. The inference fits the data and predicts a testable case, which is a brand whose citation share rises in a distant subject while its mention share does not move.
Neither measured nor inferred: that moving into an adjacent category causes the mention share to rise. No study has assigned coverage rather than observed it, none has controlled for the writing quality, brand size or third-party activity that plausibly move both, and none has been replicated by a second party. The pattern is strong enough to sequence work against and not strong enough to promise anything.
Frequently asked questions
How do I decide which category to enter next?
Rank candidates by how close their buyer questions sit to the questions you already answer well, using embeddings and cosine similarity, then filter the top of that list by whether the categories share a customer, a problem, a product capability and a buying process. Proximity is the variable that tracked the outcome in the only data available.
Can I compute topical distance without a vendor tool?
Yes. It needs a list of candidate categories, roughly five representative buyer questions for each, the same for your established subjects, and any current embedding model. The scores rank candidates against each other; they are not on the same scale as any published figure.
What is a good citation-to-mention conversion rate?
The published reference points are about 46% for categories most related to a brand's core and about 18% for the least related, both from a single 2026 panel. Use them as an orientation rather than a target: a new category converting near the low end suggests the subject sits further from your core than the plan assumed.
Why do my citations go up while brand mentions stay flat?
That is the expected pattern when publishing into subjects distant from your established ones. Citations respond to documents matching queries. Named mentions reflect an association between your brand and a subject, which your own publishing influences far less directly.
Does this apply outside ChatGPT?
Unknown. Every figure on this page comes from a panel that measured ChatGPT in the United States over six months. Gemini, Perplexity, Copilot and Google's AI Overviews were not included, and no equivalent measurement of them has been published.
Is entering a new category with one article actually harmful?
The panel found single-prompt presence associated with a lower mention share, with the effect clearing at three of five prompts. Whether shallow coverage causes that or simply describes what an unfamiliar brand looks like in a category is not something the study design can separate.