AI Search Visibility Metrics & KPIs: a real five-week case study
Most AI visibility reports stop at the dashboard. Ask what the real AI search visibility metrics and KPIs should be, and most platforms give you the same five terms — visibility, share of voice, sentiment, position, citations — and stop there. This one didn't, and that's the only reason the story below got caught at all.
By Philipp Enders·Founder, CrunchJunkieBuilds the reporting and AI-visibility tooling this analysis was run with.
The setup: five weeks of rising AI visibility and share of voice
A luxury property brand came into this tracking dominant. Across 145–159 scan runs a week — spanning six of CrunchJunkie's eight tracked AI models, including Claude, ChatGPT, Perplexity, Gemini, Google AI Overviews and Google AI Mode, not a single lucky prompt — its AI Visibility Score climbed steadily over five weeks: 31%, 32%, 33%, 35%, 35%. That climb came with a real margin of error attached (±3.8 percentage points), the kind of detail that separates a genuine trend from a coincidence of small sample sizes.
Share of Voice told the same story, louder. Against every competitor tracked, this brand held 87% share of voice. The next-best competitor: 5.9%. That's not a close race — that's a brand that's essentially the only name AI models reach for in its category. It's the difference between winning the room and merely being in it.
By every visibility metric on the market, this was a five-week win streak.
The catch: why Average Position contradicted Share of Voice
Being mentioned isn't the same as being recommended, and Average Position is where that distinction shows up. Despite the dominant visibility and share of voice, this brand's average position across scan runs was 2.8 — while a much smaller competitor, appearing in a fraction of the mentions, averaged 2.0 on the rare occasions it did show up. Dominant presence and prominent placement turned out to be two different wins, and this brand had won only one of them.
Sentiment added a similar wrinkle. The brand scored 72/100, ahead of a competitor sitting at 50/100 — a clear lead, until you look at where that competitor's score came from: two or three mentions per model, nowhere near enough volume to trust. A sentiment score built on 45 mentions carries real signal. One built on a handful doesn't. Without sample sizes attached, that "loss" on paper would have looked meaningful. With them, it barely registers.
Citation tracking surfaced the most actionable gap of all. The brand's own site appeared as a retrieved source in 30% of scan runs — AI models were actively pulling from its pages nearly a third of the time. That's real authority already in the pipeline. But appearing as source material and being formally cited turned out to be two different stages, and the brand hadn't yet closed that gap — a clear next step (stronger structured data, more quotable content) rather than a mystery.
The number that mattered: AI visibility up, website sessions down
None of the above would have raised any alarms. Rising visibility, dominant share of voice, a sentiment lead, real citation potential — read in isolation, every AI-visibility metric on the board said things were going well.
Then the website numbers came in for the same five weeks. Sessions fell every single week: 458, 425, 380, 300, 245.
Visibility went up. Traffic went down — steadily, over the exact same five weeks the AI metrics were climbing.
AI visibility
31% → 35%
Website sessions
458 → 245
Same client, same five weeks. Visibility climbing, sessions falling — a gap only visible when both are on one page.
Why an AI visibility dashboard alone can't catch this
This wasn't a bad week for the brand's marketing team, and it isn't a reason to distrust the visibility numbers — both things were true at once. It's the reason AI visibility can't live in its own dashboard. A visibility-only view would have shown five straight weeks of good news. Only by putting the AI metrics and the GA4 numbers in the same report did the real question surface: why is a brand becoming more visible in AI search and less visible to its own website at the same time?
That's a question worth investigating — a channel shift, a seasonal dip, a landing page problem, an attribution gap where AI-driven visits are landing as direct or organic traffic instead of being credited to AI. It might be all four. But you can't ask the question until you can see both trends on one page.
The takeaway: tie AI visibility metrics to conversions
Every metric in this story — visibility, share of voice, position, sentiment, citations — was real, correctly measured, and individually good news. None of them, alone or together, would have caught the traffic decline running underneath them.
That's the case for tying AI visibility to conversions: connecting AI visibility to the outcomes it's supposed to predict. Not because the visibility metrics are wrong, but because they were never designed to tell you whether they're working.
See how CrunchJunkie connects your AI visibility to GA4.
Frequently asked questions
Visibility score, share of voice, average position, sentiment and citations are the five most platforms report. All five are real metrics, but they measure presence in AI answers, not the outcome of that presence. In the five-week case study above, every one of them improved while website sessions fell 458 to 245 over the same period. The metric that mattered was the one no visibility dashboard holds on its own: traffic and conversions from the same weeks, read on the same page.
There are four common explanations and they are not mutually exclusive: a channel shift, where AI answers satisfy the question without a click; a seasonal dip unrelated to AI; a landing-page or conversion problem downstream; or an attribution gap, where AI-driven visits arrive tagged as direct or organic and are never credited to AI at all. Distinguishing between them requires AI visibility data and analytics data in one view — a visibility-only dashboard cannot separate them, because it cannot see the traffic side.
Enough that a single answer cannot move it. In the case study, a brand scoring 72/100 across 45 mentions was compared with a competitor scoring 50/100 across two or three mentions per model. The second number looks like a result and is closer to noise. This is why sentiment should never be shown without its mention count attached: the score alone gives a small sample the same visual weight as a large one.
A retrieved source is a page an AI model pulled from while composing its answer. A citation is a formal, visible reference to that page in the answer itself. They are separate stages, and the gap between them is actionable. The brand in this case study appeared as a retrieved source in 30% of scan runs without a matching citation rate — meaning the authority already existed and the next step was making the content easier to quote and attribute, through stronger structured data and more self-contained, quotable passages.
Not on its own. It can tell you whether a brand is present, prominent, and positively described in AI answers, which is genuinely useful. It cannot tell you whether that presence produced sessions, leads or revenue, because those numbers live in analytics. Answering the question the client actually asks — is this working — requires both sets of numbers in the same report, covering the same weeks.
See your AI visibility on your own brand
Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.