Published
The decision this is actually about
Almost no content plan is a choice between publishing and not publishing. It is a choice among subjects that all look reasonable, made under a budget that will not cover all of them, and the tiebreaker is usually estimated search volume.
Volume is a measure of how large a category is. It carries no information about whether the category is winnable by you, and for AI search visibility there was until recently nothing that did. The 2026 work on ChatGPT category prompt sets supplies a candidate variable, and the useful thing about it is that a practitioner can compute it without buying anything.
What follows is a planning sequence built on that. It is worth stating up front what kind of evidence sits underneath: cross-sectional observation of one vendor's panel, no controls, no replication. That is enough to sequence work against. It is not enough to promise an outcome, and any version of this method presented as a formula is overselling what was measured.
Step one: find your anchors, and expect fewer than you think
An anchor is a subject where you already earn both kinds of visibility across most of the questions buyers ask inside it: your pages get cited as sources, and your brand gets named in the answer text. Both, not either.
The distinction is the entire point of the exercise. A subject where you are cited often and never named is not an anchor. It is a subject where your documents perform and your standing does not, and expansion outward from it is expansion from a position you do not hold.
Build the list by taking each subject you believe you own, writing out five representative buyer questions — the definition question, the comparison, the alternatives, the fit question, the purchase question — and checking both signals across all five. Any tool that separates citations from mentions at the category level will do this. So, more slowly, will running the prompts by hand and recording what comes back, which has the advantage of showing you the actual answers.
Most organizations finish this step with one or two anchors and a longer list of subjects they had assumed were anchors. That correction is worth the exercise on its own.
Sampling matters more here than it looks. These systems do not return the same answer to the same question twice, so a single run of five prompts is not a measurement, it is one draw. Repeat each prompt several times across separate sessions before recording a result, and record how often the brand appeared rather than whether it appeared. A brand showing up in three runs out of ten is a different finding from one showing up in nine, and a single pass cannot tell them apart.
Step two: rank candidates by distance, not by volume
The variable that tracked category outcomes in the published panel was proximity: how closely the questions in a new subject resemble the questions in subjects the brand already answered well.
Computing it is ordinary work. Write five buyer questions for each candidate subject. Embed those question sets, and the question sets for your anchors, with any current text embedding model. Take the cosine similarity of each candidate against its nearest anchor. Sort descending.
The output is a ranked expansion list grounded in the same construct the research used, and it will disagree with your keyword tool's ranking. That disagreement is the value; if the two agreed you would not need the exercise.
Two corrections belong on top of the raw scores. Similarity computed on question text is measuring vocabulary, which is a proxy for market adjacency rather than the thing itself — two subjects can share language and share no customer. Filter the top of the list by whether a candidate shares a customer, a problem, a product capability and a buying process with the anchor it scored against. Where the score is high and those four are not, the score found words.
The second correction is that your scores are not on the same scale as any published figure. They rank candidates against each other. They do not predict a result.
A worked example makes the filter concrete. Suppose the anchor is payroll software. Small business banking scores high on question similarity and passes the commercial filter as well: same customer, adjacent problem, plausible product overlap, comparable buying process. Human resources compliance also scores high and passes on customer and problem while failing on product, which makes it a candidate for content and not for a product claim. Warehouse robotics scores low and is correctly excluded. The instructive case is a subject like employee engagement software, which shares vocabulary heavily with payroll content, shares a buyer, and involves an entirely different purchase — high score, partial pass, and a decision that needs a person rather than a threshold.
Step three: enter at depth, because partial entry is not a cheap test
Publishing one piece into a new subject to see whether it gains traction has been standard practice for as long as organic search has existed, and it has been sound practice, because a page that fails to rank simply sits there costing nothing else.
The measured panel puts a question mark over that. Brands present in only one of a category's five prompts were associated with a lower mention share, and the effect did not clear until three of five. Partial presence looked worse than absence.
Whether that is causal is genuinely unsettled — a brand in one prompt out of five is also, plausibly, just a brand nobody has heard of — so the honest way to use it is as a planning floor rather than as a warning about harm. Treat three of five buyer questions as the minimum unit of category entry. The consequence is a scheduling constraint more than a spending one: the same budget buys three subjects covered properly or twelve covered shallowly, and on this evidence those two plans are not equivalent.
Cover the sequence rather than the volume. A definition page that ranks does nothing for the comparison prompt or the alternatives prompt, and those are the questions where a buyer is choosing.
Step four: fund the signal that actually gates expansion
Two signals behave differently, and most programs fund only one of them.
Citations respond to publishing. A document matched a query and cleared the source conditions, and more documents produce more chances of that happening — including in subjects where a brand has no standing at all, which is why citation counts travel across category boundaries so easily.
Named mentions respond to something else. A model names a brand because its representation associates that entity with that subject, and that association was built from the corpus rather than from any single retrieval. The corpus is mostly other people's writing: review platforms, industry publications, community threads, expert commentary.
A program that funds content production alone is funding the signal that spreads and not the one that gates. The inputs on the mention side sit in different parts of most organizations — product marketing owns review platform presence, communications owns trade coverage, whoever owns the customer relationship owns community reputation, and entity consistency across schema, naming and third-party profiles sits with technical SEO. Assigning the whole outcome to one team and asking for AI visibility as a deliverable gives that team perhaps a third of the levers.
Step five: measure the ratio, not the counts
The reporting most teams have shows citation count and mention count as two totals, or worse, one composite score built from both. Neither answers the question a plan like this raises, which is whether a subject you entered is one you belong in.
The measurement that does is the share of your citations in a category that also carry a named mention. It is insensitive to publishing volume, which is exactly why it works as a check: you can lift a citation count indefinitely by publishing more, and the ratio will not move unless recognition moves with it.
The published reference points come from the one panel that has measured them — roughly 46% in the categories most related to a brand's core, roughly 18% in the least related. A new subject sitting near the low end is telling you the content arrived somewhere your standing has not followed. That is not a reason to abandon it. It is a reason to check whether the adjacency filter in step two was right, and to look at what the mention side is being funded with.
Watch the margin as well as the level. Leads under a couple of percentage points in mention share flipped month to month in the panel; leads near three points held. Below that, a monthly report is describing variance.
What a monthly report should carry, then, is four things per subject: the citation count, the named-mention count, the ratio between them, and the point margin against the nearest competitor on mention share. Four numbers, per category, with no composite score built on top of them. A single visibility figure aggregates a signal that travels with a signal that gates, weights them by whichever is more numerous, and hides the only comparison that answers the question the plan was built to answer.
What this method does not do
It does not establish that any of this causes anything. The underlying work is observational, and its authors were careful to write associated with. A variable that accompanies visibility may produce it, follow from it, or share a cause with it, and no published study has assigned coverage rather than observed it.
It does not transfer across surfaces on any evidence. Every figure here comes from a panel measuring ChatGPT in the United States over six months. Gemini, Perplexity, Copilot and Google's AI Overviews were not measured, and assuming they behave alike is an assumption, not a finding.
It does not help in every sector equally. In the measured panel, legal and healthcare showed a ceiling that coverage depth did not lift. If you work in those fields, treat depth as the entry requirement and put the marginal budget on credentialing, original data and independent validation — while noting that this is a hypothesis about a ceiling nobody has explained.
What the method does do is replace a tiebreaker that was never connected to the outcome with one that at least tracks it in the only data anyone has published. That is a modest claim, and it is the accurate one.
Frequently asked questions
How many subjects should a plan cover at once?
Fewer than most plans do. The measured panel found breadth without depth to be the configuration with the weakest outcome, and put the floor for meaningful presence at three of a category's five buyer questions. A budget spread across a dozen subjects at one asset each buys a position in none of them.
Do I need a vendor tool to run this?
Not for the adjacency ranking, which needs candidate subjects, five buyer questions each, and any embedding model. You do need something that separates citations from named mentions by category to measure the result, and that is either a tracking tool or a disciplined manual sampling routine.
What if my anchors turn out to be subjects where I am cited but never named?
Then they are not anchors, and expansion from them starts from a position you do not hold. The work in that case is on the mention side within the subject you thought you owned, rather than on entering a new one.
How long before a new category shows anything?
Unknown, and nobody has published an answer. The panel this method comes from observed categories monthly for six months without running interventions, so it has nothing to say about how quickly a change in coverage registers, or whether it registers at all.
Is search volume useless for this?
No. It tells you the size of the prize, which matters once you have a shortlist you can plausibly win. The argument here is that it is the wrong variable to sort by first, since a large subject far from your standing converted citations into mentions at well under half the rate of a smaller one nearby.