rec 20 · re-instrumented 2026

first published 2026-08-09

The prompt is not the query. Stop optimising the wrong string.

The prompt-screenshot slide is in every SEO deck: fire prompts, count brand mentions, call it visibility. Why prompt tracking measures the wrong thing — and what the models actually consult.

2,161 words · 10 min read · 14 min listen

read by jamie mckaye — his own voice, via his voice model. not a studio take.

00:00 / --:--

Every SEO deck I've seen in the last six months has the same slide. It shows a list of prompts — "best CRM for small teams," "top project management tools for agencies," "which analytics platform for e-commerce" — and next to each one, a screenshot of ChatGPT, Perplexity, or AI Overviews with the client's brand either present or absent. The exercise is called AI visibility tracking, or GEO auditing, or prompt monitoring depending on who's selling it. The methodology is identical everywhere.

It's also mostly wrong.

Not wrong in the sense of "the tool is bad" or "the tracking is inaccurate," though both are often true. Wrong in the sense that the thing being measured is not the thing that decides whether a real buyer ends up thinking about your brand. The category prompt is the SEO industry's default because it looks like a keyword. It behaves nothing like one, and the buyers using AI systems are not typing it.

Slobodan Manic made this exact point in Search Engine Journal this week, tucked into a broader argument about the GEO naming fight. The line worth pulling out: "Most people testing AI visibility type the category prompt, 'best CRM for small teams,' and check." Reverse-engineer what the model already believes about you, he argues, not what it says when prompted with a shape that mimics a Google query.

He's right, and the implication is bigger than the article makes it. The prompt-tracking industry has spent eighteen months building tools that measure the wrong string. The right unit of measurement isn't the prompt. It's the model's underlying belief about your brand — and those two things are related, but they are not the same.

The category prompt is a keyword in cosplay

Here's what's happening when someone sells you a prompt-tracking dashboard. They pick a list of category-defining questions that look like commercial-intent keywords. They run them against three or four LLMs on a schedule. They tell you whether you appeared, ranked, or got skipped. They chart it over time. Some of them scrape the linked citations too.

The category prompt is what an SEO would type. That's not a coincidence. It's what the tools measure because it's what the people building the tools recognise as a keyword.

The whole exercise is rank tracking with a chatbot skin. And the reason it feels satisfying is that it produces a scoreboard-shaped output — line goes up, line goes down, competitor overtakes you — which is exactly what SEOs are trained to interpret.

The problem is that category prompts are not what buyers actually type. Buyers ask messy, conversational, context-loaded questions. "We're a fifteen-person agency, our current PM tool is Asana but the reporting is killing us, we bill hourly and need to track profitability per client — what would you switch to?" That prompt does not appear on any tracking dashboard, because no vendor can enumerate the infinite variations of it. But that's the shape of the actual query.

The category prompt is what an SEO would type. That's not a coincidence. It's what the tools measure because it's what the people building the tools recognise as a keyword.

Kevin Indig's H1 2026 halftime report has a data point that should be on every consultancy's wall: 91% of citations appear in only one of ChatGPT, Perplexity, or AI Overviews — never in more than one. Same query, three systems, three different answers. The variance across the same prompt is enormous, and the variance across similar-but-not-identical prompts is enormous inside a single system too.

What this means, if you take it seriously, is that no single prompt result tells you anything durable about your position. You could win the "best CRM for small teams" prompt in ChatGPT this week and get skipped by it next week — same prompt, same account, same model version — because a marginal shift in retrieval or reasoning surfaces a different set of candidates. Prompt tracking treats these results as measurements. They're closer to samples from a probability distribution, and the distribution itself is what you actually care about.

The distribution is a function of what the model believes about your brand — a composite formed from training data, retrieval-augmented sources, and whatever real-time web fetching the system does. Indig makes the point in the same report that brand mentions correlate more closely with real business outcomes than citations do, precisely because mentions are a proxy for the underlying belief, and citations are a downstream artefact of it. Chase the artefact and you'll spend a lot of money moving a number that doesn't correlate with pipeline. Chase the belief and the artefact takes care of itself.

The prompt is the query. The belief is the target. The industry is optimising for the wrong string.

What "the belief" actually is, in mechanical terms

Let me be specific about what I mean, because "the model's belief about your brand" sounds mystical and it isn't.

When an LLM answers a category question, it's doing two things in sequence. First, it's assembling a set of candidate entities that fit the semantic shape of the query — for "best CRM for small teams," that's a shortlist of CRMs the model associates with the small-team context. Second, it's ranking those candidates by whatever internal weights it's built up from training and retrieval about credibility, popularity, fit, and specificity to the modifier.

The prompt only influences step two. Step one is entirely a function of what the model already thinks exists in the world and which entities belong in this category. If your brand doesn't make the shortlist, no amount of prompt-level optimisation gets you into the answer. You could have the world's most perfectly structured comparison page and it wouldn't matter, because the model never considered you as a candidate.

This is why Manic's advice to prompt the LLMs about your brand and your products — not the category — is the actual move. When you ask "what is [your brand] known for" or "who uses [your product]" or "compare [your brand] to [competitor]," you're not measuring your rank. You're reading out the model's stored representation of you. That representation is what determines whether you get shortlisted for the messy, context-loaded queries buyers actually type.

If the model thinks you're a mid-market CRM for professional services firms, you'll surface in prompts shaped like that context, regardless of exact wording. If the model has a confused or thin representation of you, you'll surface intermittently in ways that look like noise on a prompt-tracker but are actually the model guessing.

Why the industry ended up here

The prompt-tracking industry didn't happen by accident. It happened because it looks familiar. SEOs know how to track keywords. Rank tracking is the default deliverable in every agency retainer. When AI answers arrived and clients asked "how do we know if it's working," the fastest path to a billable answer was to reskin rank tracking with a chatbot backend and call it GEO monitoring.

I don't blame the vendors for building it. There was demand, they filled it, the dashboards look nice. What I mind is the framing — the pretence that prompt-level tracking is a rigorous methodology rather than a legacy shape being forced onto a fundamentally different mechanism. Indig's own suggestion in the halftime report is that prompt tracking should behave more like polling and focus-group research than like rank tracking, and that's the framing I'd adopt if I were building tooling in this space today. You'd sample a distribution across query variants, weight by realistic usage patterns, and report on confidence intervals rather than binary pass/fail.

Nobody's doing that yet because it doesn't look like an SEO deliverable. It looks like a research report. And research reports are harder to sell on a monthly retainer than a dashboard with green and red arrows.


What to measure instead, without pretending you have perfect tools

I want to be honest about the limits of the alternative, because the flip side of "prompt tracking is the wrong measure" isn't "here's the right one." The right measure doesn't exist as a productised tool yet. But there are three things I'd be doing with clients right now that are more informative than category-prompt monitoring.

Read out the model directly. Prompt LLMs about your brand, not your category. Ask what you're known for, who uses you, what you're compared to, what your reputation is. Do this across models. Compare the answers. The gap between what the model says and what you actually are is your discovery problem stated in mechanical terms.

Track brand mentions, not citations. Indig's data point about mentions correlating with outcomes is worth taking seriously. Mentions across podcasts, YouTube transcripts, Reddit threads, industry press, and independent blogs are what train the model's belief in the first place. If you can measure the volume, sentiment, and consistency of those mentions, you're measuring the actual input variable, not the downstream artefact.

Audit for hallucination surface area. Manic's specific suggestion — check whether what you're giving the models would cause them to hallucinate — is the most operationally useful piece of advice I've read on this topic all year. Every key page on your site, every off-site profile, every third-party listing should say a consistent thing about who you are. Inconsistency is what produces the confused, thin brand representation that surfaces as flaky prompt results. Consistency across every source a model can read is how you make yourself easy to describe.

None of these are dashboards. They're research activities. They take time, they require judgement, they don't produce a single number the client can screenshot for their board deck. That's exactly why they're worth doing — the vendors haven't productised them yet because they resist productisation.

The naming problem in miniature

There's a broader point sitting under all of this that connects to the GEO naming fight I wrote about last week. The prompt-tracking industry is a specific instance of the general disease: taking a legacy SEO shape and forcing it onto a mechanism that doesn't have the same substrate underneath.

Rank tracking worked in classic search because the ranking function was deterministic-ish, the query surface was finite-ish, and the same query produced the same result across users within reasonable bounds. None of those conditions hold for LLM answers. The output is probabilistic, the query surface is effectively infinite, and the same prompt produces different answers across users, sessions, and days.

Forcing rank-tracking methodology onto that substrate isn't just imprecise — it's category error. You're using a ruler to measure temperature. The ruler will produce a number. The number will not tell you anything true about how hot it is.

The reason this matters isn't philosophical. It's that every hour spent chasing prompt-level rankings is an hour not spent on the underlying belief work — the brand consistency, the entity clarity, the earned mentions across places the model reads. And that underlying work is what actually moves the outcome, whether the answer surface is a Google result, a ChatGPT reply, a Perplexity citation, or whatever comes next.

Where I'd concede

The honest limit here is that measurement matters, even when the thing you're measuring is imperfect. Clients need to see something. Boards need to see something. Consultants who bill for retainers need to demonstrate progress in a form that doesn't require a philosophy lecture to interpret.

I don't think the answer is to abandon prompt tracking entirely. I think it's to be honest about what it measures and what it doesn't. If you want to run category prompts on a schedule as a rough temperature check, fine — treat them as one input among several, expect volatility, don't build strategy on top of them. The moment you start optimising specific pages to move a specific prompt result, you've drifted into cargo-cult territory. You're doing 2015 SEO on a system that doesn't have a link graph.

The people I trust in this space — Manic, Indig, Alderson, Oberstein — are all pointing at variations of the same underlying truth. The prompt is not the query. The tracker is not the outcome. The belief is the target, and the belief is built by the unglamorous, slow, hard-to-productise work of being coherent about who you are across every source a model can find.

That work isn't new. It isn't GEO. It isn't AEO. It's what good SEO always meant, before the industry decided it was easier to sell dashboards than substance. The tools have caught up to the wrong methodology first, which is depressingly on-brand for our industry. Give it a year and someone will build the polling-style measurement Indig described. Until then, the honest answer to "how do we know if it's working" is a research report, not a screenshot.

The category prompt on your tracking dashboard is a keyword pretending to be a query. Stop optimising for the pretend one. Optimise for the belief that produces both.

self-audit

agent-ready grader

live

run it at /lab

mcp tools

04

public — /api/mcp/

corpus

222→0

retired into the field

build

p5

hardening pass

your visit — measured on you, just now

ttfblcpinpawaiting inputcls

compiled from c2660ba · 2026-08-26 16:14 utc · push = ship


Jamie McKaye — technical SEO, AI systems, full-stack build, technical writing. One person, no handoffs.