Category coding is the SEO story of the year
New research shows LLMs recommend brands based on category coding, not entity strength. It rewrites most of the current GEO playbook.
There's a study out this week from João da Silva and Andrea Volpini that should be the most-discussed piece of research in the industry right now. It isn't. It sits behind a Search Engine Land byline, gets a polite nod from the entity SEO crowd, and then the news cycle moves on to whatever Google shipped in Merchant Center this afternoon.
That's a mistake. The study quietly demolishes one of the load-bearing assumptions of the "get cited by AI" playbook, and almost everyone selling GEO services is going to have to reckon with it eventually.
The short version: recognition and recommendation are not the same thing. A brand can be perfectly known to an LLM and still be functionally invisible in the queries that matter. What determines whether you show up isn't how strong your entity is. It's whether the category the model has coded you into matches the category your customer is searching from.
If that sounds like a small distinction, stay with me. It rearranges most of what the industry has been saying for the last eighteen months.
What the study actually measured
Da Silva and Volpini ran 14,140 API calls across ChatGPT, Gemini, Perplexity, Claude and Google AI Overviews, testing twelve UK athletic apparel brands under two different category framings: "athleisure" and "athletic footwear." Same brands, same models, same week. The only variable was the category word in the prompt.
The results were symmetric in a way that rules out noise. New Balance appeared in 1% of athleisure prompts and 90% of footwear prompts. Lululemon went the other way — 90% of athleisure prompts, 0% of footwear. Alo Yoga: 63% and 0%. Gymshark: 37% and 0%. The Knowledge Graph score of the brand — the industry's favourite proxy for entity strength — barely predicted the outcome. New Balance has a KG score of 64,235. Lululemon has 810. Both are strong entities. Both are recognised universally. Neither of those facts was relevant to the actual recommendation behaviour.
What predicted the outcome was what the authors call category coding. Two ingredients: the Knowledge Graph description field, and the third-party corpus of editorial content that has accumulated around the brand in the context of a specific category. Nike, New Balance and Reebok all share the same KG description — "footwear company." So all three get recognised. But their footwear-context editorial corpus is what makes them recommendable when the query uses the word "footwear." Lululemon's corpus is built from fashion editorial, wellness roundups, and activewear listicles. When "athleisure" is the frame, that corpus is where the model looks. When "footwear" is the frame, that corpus is invisible.
Why this breaks half of the current GEO advice
The dominant GEO framing right now — the one being sold in strategy decks across the industry — treats AI visibility as an entity problem. Build the Knowledge Graph. Add schema. Get more press. Strengthen your entity so the LLM recommends you.
Recognition is table stakes. Category coding is the moat.
That framing assumes the model evaluates your brand against a query and decides whether you're good enough. What actually happens is the reverse. The model evaluates the query, extracts the category signal, and looks for brands coded into that category. Your entity strength is a filter, not a driver. If you're not category-coded into the framing the user is using, no amount of entity signal will pull you in.
This is why so many "well-optimised" brands are quietly missing from AI results in categories they clearly compete in. The client isn't invisible because their entity is weak. They're invisible because they've been coded into the wrong bucket by the third-party corpus.
Recognition is table stakes. Category coding is the moat.
I've been describing this to clients for months without having the terminology for it. A UK homeware brand I've worked with is a good example of the shape — they'd been on every "best of British design" roundup you can name. Universally recognised. Zero AI citations for the actual purchase-intent queries their customers use, because the editorial corpus around them was all lifestyle-magazine coverage, not category-comparison coverage. The model knew who they were. It just didn't know they were an answer to the questions being asked.
The measurement problem this creates
Here's where it gets uncomfortable for the prompt-tracking industry. If category framing changes the citation set by 89 percentage points on the same brand in the same week, then the current generation of AI visibility tools — the ones that pick a prompt set and track which brands appear — are measuring the category framing the tool vendor chose, not the visibility of your brand.

Every one of those dashboards is running a fixed prompt list. The choice of prompt IS the measurement. If your tool tests "best athletic footwear brands" and your competitor's tool tests "best athleisure brands," you'll get diametrically opposite readings of the same market. Neither is wrong. Neither is complete. Both are being sold as ground truth.
I've written about this measurement gap in the context of prompt-tracking tools more broadly, but the da Silva/Volpini data adds a specific mechanism to what was previously a general suspicion. It's not just that prompt-tracking is noisy. It's that the prompt is the variable, and the tools treat it as a constant.
The honest version of an AI visibility audit has to include category framing as a dimension. You need to know which frames your brand appears in, which frames it doesn't, and — this is the strategic question — which frames your actual customers are using.
The fix isn't more entity work
If your brand is category-miscoded, more schema won't fix it. More press coverage won't fix it, unless the press coverage is specifically inside the category corpus you want to be in. Getting into another "innovative UK brands" listicle strengthens your entity but does nothing for category coding. Getting into a specific comparison piece — "best athleisure brands for pilates" or whatever the equivalent is in your vertical — changes the corpus signal in a way that entity work cannot.
This reframes what earned media is actually for in an AI search context. It's not authority-building in the abstract. It's category-encoding in the specific. Every piece of coverage either reinforces the category framing the model already has for you, or it doesn't. Coverage that doesn't slot into the category you want to be discoverable in is, from an AI recommendation standpoint, close to worthless.
That's a genuinely different way of thinking about PR strategy. Most in-house teams and most agencies are still optimising for volume and prestige of coverage. The da Silva/Volpini finding suggests the metric that matters is category alignment. A single Wirecutter comparison in your target category is worth more than a hundred general "brands to watch" mentions.
What this means for the "brand is the moat" argument
Regular readers know I've been banging the brand drum for a while. Brand recognition genuinely does drive AI citations, and building brand is the highest-leverage discovery activity for most businesses. That case still holds. But this study forces a refinement: brand is the moat when your brand is coded into the right category. Brand equity that lives in the wrong category is stranded equity.
Think about what this means for a company doing a category expansion. If you've built brand strength in one category and want to be discoverable in an adjacent one, the AI systems will treat you as if you don't exist in the new category until the third-party corpus catches up. Your entity strength doesn't transfer. Your recognition doesn't transfer. You essentially have to earn category coding from scratch, and the mechanism for earning it is editorial coverage inside the new category — not more schema, not more of the coverage that got you into the first category.
This has real strategic implications. It means category expansion in an AI-mediated discovery world is materially harder than it was when Google's classic algorithms were more forgiving of adjacent-category authority transfer. It also means brands with narrow, deep category coding are more defensible than brands with broad but shallow coding across many categories.
The honest limits
The study is on one vertical — athletic apparel — and one geography, the UK. Category coding could work differently in categories with less editorial density, or in B2B categories where the corpus is more thinly distributed across trade publications and Reddit rather than mainstream press. We don't know yet. The mechanism is likely to generalise, but the specific numbers won't.
The other limit is time. Category coding isn't static. As the third-party corpus around a brand changes, the coding drifts. New Balance's footwear coding is stable because decades of running-shoe editorial reinforce it. A newer DTC brand's coding is far more volatile — one influential comparison piece can move it. That's a threat if you're miscoded, and an opportunity if you're not yet coded at all.
And there's the obvious caveat that the underlying models are changing. What Claude does with category signals today isn't necessarily what Claude 5 does next quarter. The mechanism is likely durable — it's not a bug, it's how retrieval-augmented generation works with editorial corpora — but the specific behaviours will shift.
What to actually do about it
Three things, in order of leverage.
First, audit your category coding before you audit anything else. Take the five to ten queries your actual customers use — the language they use, not the language your marketing team uses — and check whether your brand appears in AI responses to those queries across ChatGPT, Gemini, Perplexity and AI Overviews. If you don't appear, check whether you appear when the category term is changed. That tells you whether you have a coding problem or a broader authority problem. They require different fixes.
Second, look at where you actually get press. Not the volume of it, not the prestige of the mastheads, but the category context of each piece. Are you being covered inside the categories your customers search from, or adjacent to them? Adjacent doesn't count. This is a genuinely different way of briefing PR agencies, and most of them will resist it because "get us into any tier-one publication" is easier to sell than "get us into three specific comparison pieces in this specific category."
Third — and this is the one nobody wants to hear — reconsider what your entity SEO programme is actually optimising for. If it's producing more Knowledge Graph strength without moving category coding, it's producing recognition without recommendation. That's not nothing, but it's not what most clients think they're paying for.
The industry has spent two years telling businesses that AI visibility is downstream of entity strength. The da Silva/Volpini data suggests it's downstream of category coding, and entity strength is one input among several. That's a smaller, sharper claim than "build a strong entity," and it points at very different work.
Most GEO strategy decks being shown to UK businesses right now are still operating on the older, wrong assumption. Category coding is the SEO story of the year. It deserves more than one Search Engine Land byline and a shrug.
Ready to improve your visibility in AI search?
If you're an SME in Surrey or London and you want more qualified leads from search — including the growing AI answer layer — let's talk.
Book a discovery call