All posts
AI visibility·23 August 2026·8 min read

Ask Five AI Engines About Your Brand, Get Five Answers — Here's What to Measure Instead

Ask ChatGPT which agency to hire and it names you. Ask Gemini the same question an hour later and your name is nowhere. Neither engine is broken, and neither is lying — they are two different systems reading two different slices of the web, and each one is also a little different from itself between one answer and the next. That is the uncomfortable truth under every AI-visibility dashboard: there is no single, stable “what AI says about you.” There is a distribution, and the only honest way to report it is to sample it properly and show the uncertainty. Here is why one check misleads you, and what to measure instead.

By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.
One prompt on the left fanning out into five different answers across five AI engines — the same question producing five different results.

Do AI engines really disagree about the same brand?

Yes, and the gap is wider than most people expect. In August 2026 a developer ran the same twenty queries across five engines — ChatGPT, Claude, Gemini, Perplexity and Bing Copilot — against two of his own sites. Only two engines ever cited him at all, and the two never once agreed on the same page. As he put it, "each one only saw a fifth of the picture" (Nwaneri, "I Tested 5 AI Engines On My Own Sites. None Agreed."). It is one site and one snapshot, so read it as a lived example rather than a benchmark — but it is the shape everyone who runs this test finds. The larger studies point the same way. Vendor analyses of which sources each engine cites put the overlap remarkably low: one Q2-2026 benchmark measured a mean cross-engine citation overlap of 0.18 on a Jaccard scale, and a separate analysis found only about 14% of top cited domains shared between engines. Treat those exact figures as directional vendor analysis, not gospel — but the direction is unmistakable, and even Google's own two AI surfaces, AI Overviews and AI Mode, reportedly share only a small fraction of their citations. Different engines draw from different indexes, weight different signals, and answer at different moments — so "our AI visibility" was never one number. It is at least one number per engine. The practical consequence is blunt: check a single engine and you have roughly a one-in-five chance of drawing the exact opposite conclusion to the one the engine beside it would give you.
SourceThe testWhat it found
Thinking Machines LabSame prompt, same model, 1,000 runs at temperature 080 different answers
Microsoft (Phi-4 report)Same evaluation, 50 runs per model30–70 pt accuracy swings; “a single run can mislead”
dev.to (5-engine test)One brand, 5 engines, 20 queriesOnly 2 engines cited it — zero overlap
Vendor analysis (Foglift, Frase)Which sources 4–5 engines cite~14–18% domain overlap

Isn't checking once enough?

Even within one engine, one check is not enough — because the same model does not give the same answer twice. The cleanest demonstration comes from Thinking Machines Lab, who asked one model the same question a thousand times at temperature zero, the setting that is supposed to make output deterministic. They got eighty different answers. Every completion was identical for the first hundred-odd words, then diverged (Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference"). If a model can't reproduce itself a thousand times over on a trivia question, a single scan of your brand is a coin flip you happened to catch mid-air. That variance is large enough to flip conclusions. Microsoft's Phi-4 reasoning report ran the same evaluation fifty times per model and watched accuracy swing across enormous ranges — tens of percentage points for the same model on the same test (Microsoft, Phi-4-reasoning Technical Report). Swap "models" for "brands" and you have the case for repeated sampling in one sentence. A visibility score built on one run per prompt is measuring noise and calling it signal. This is why CrunchJunkie samples each prompt several times — three to ten, depending on plan — rather than once. More runs is not thoroughness for its own sake; it is the only way to tell a real change from the model clearing its throat.
Any comparison among models using a single run can easily produce misleading conclusions.
Microsoft, Phi-4-reasoning Technical Report (2026)

What a single number hides

"You're at 34% visibility" feels like a fact. On its own it is barely half of one. Two brands can both sit at 34% where one figure rests on two hundred runs that mostly agree and the other on twelve runs that scatter from 0 to 80% — completely different situations that a lone percentage flattens into one. The fix is old and boring and correct: report the sample size and a margin of error next to every rate, and never quote one without the other. CrunchJunkie shows "based on N runs" and a boundary-safe interval — one that can never imply a visibility below 0% or above 100%, and that widens when the engines disagree or the sample is thin. When that band is wide, the honest reading is "we don't know yet," not "34%." A number that cannot be wrong is a number that is not telling you anything. There is a second trap hiding in the average. Put German and English prompts in one project and a blended visibility figure describes neither audience — it merges two different markets into a number that answers nobody's question. Rates get pooled within a market, never averaged across markets. Keep the populations apart or the headline is fiction.
During development of the tool, we ran it on our own agency first — and that's where the product's core decision came from. A single AI-visibility number told us almost nothing. The day we split it by engine it fell apart in the most useful way: we were winning Perplexity and ChatGPT and almost invisible on three other engines. One number had been hiding four different jobs — so we built CrunchJunkie to never report just one.
Philipp Enders, Founder, CrunchJunkie

A warning about “AI search volume”

As money moves into this category, a familiar number is appearing on dashboards: "this prompt gets 2,400 AI searches a month." Be skeptical. There is no ground-truth source for how often people ask a given question inside ChatGPT or Gemini — the providers do not publish it — so any specific monthly figure is extrapolated, usually from small browser-plugin panels skewed toward desktop Chrome users, and presented without a methodology or a confidence range. One analysis of the practice concluded prompt search volume is "not a reliable control variable," and that prompts, unlike keywords, resist volume estimation altogether (Jaeckert & O'Daniel, "Prompt Search Volume: Real Data or All Guessed?"). We deliberately do not sell that number, because we cannot measure it honestly. What we can measure is where you are actually invisible and where a win is reachable — the prompts your buyers ask that AI answers without ever mentioning you. Priorities should come from measured visibility and statistical significance, not from a volume figure with no source behind it. If a tool shows you a confident "per month" number with no methodology, ask where it came from before you plan a quarter around it.

So what should you actually measure?

Six things, and none of them is a single check. Coverage: track every engine that matters, not the one that flatters you. Because citation overlap between engines is so low, each one you skip is a blind spot, not a rounding error. CrunchJunkie tracks ten — ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, Google AI Mode, Microsoft Copilot, Grok, DeepSeek and Meta AI. Repetition: sample each prompt several times so the number reflects the distribution, not one draw from it. Uncertainty: carry a sample size and a margin of error on every rate, and read the two together. Significance: only treat a move as real when it clears a statistical bar. A drop from 34% to 31% on thin samples is usually the dice, and an alert that fires on noise trains people to ignore alerts. Separation: keep markets and languages apart, and keep the standard answer separate from the follow-up. Supply: remember the score is only the demand side. Before a model can cite you it has to reach and read your page, which your server logs record exactly — we wrote about reading those in The Supply Side of AI Visibility. A low score can be a crawler that never got in, not content that failed to impress. None of this is exotic. It is ordinary measurement discipline — sampling, intervals, significance — applied to a medium that is unusually noisy and unusually easy to fool yourself about. What makes citations move once you can see them straight is a separate, well-studied question; the Princeton "GEO" paper is the best starting point, and found the right content signals can lift visibility by up to around 40%, varying by engine (Aggarwal et al., "GEO: Generative Engine Optimization").

How to read an AI-visibility number honestly

A quick test you can apply to any dashboard, ours or anyone's. Does the figure show how many runs it is based on? Does it carry a margin of error, and does that error widen when the sample is thin? Does it keep engines and markets separate rather than blending them into one flattering average? Does a change have to be statistically significant before it is called a trend? And when it shows you a "search volume," does it tell you where that number came from? If the answers are yes, you are looking at measurement. If they are no, you are looking at a number designed to reassure you — and in a medium where one model gave eighty answers to one question, reassurance is the last thing you want. Reassurance doesn't win the citation. Measuring the right thing, repeatedly, across every engine, is what tells you where you actually stand — and what to fix first.

Frequently asked questions

Because they are different systems drawing on different indexes and weighting different signals, and they answer at different moments as the web changes. Independent tests routinely find very low overlap in which sources engines cite for the same query — often under 20% — so a result from one engine frequently does not hold on another. Tracking a brand across multiple engines is the only way to see the whole picture rather than one engine's slice.

No. Large language models are not deterministic in practice; the same prompt can return different answers on repeated runs, even at settings meant to remove randomness. One study got eighty distinct answers from a thousand identical requests. That is why a single check is unreliable and why a visibility figure should be based on several samples of each prompt, not one.

There is no single magic number, but one run per prompt is too few — it measures a single draw from a noisy distribution. Sampling each prompt several times (CrunchJunkie uses three to ten depending on plan) and reporting a margin of error lets you separate a real change from ordinary run-to-run variance. The wider the disagreement between runs, the more samples you need before a number means anything.

No tool can measure that honestly. The AI providers do not publish per-prompt query volumes, so any specific “searches per month” figure is extrapolated, usually from small, skewed browser panels, and typically shown without a methodology. Prioritise by measured visibility and statistical significance instead — where you are actually invisible and where a win is reachable — rather than by an unsourced volume number.

Often it is not a content problem at all but a supply-side one: the engine could not reach or read your page. AI crawlers get blocked by firewall and CDN rules that a robots.txt actually allows, and many live-answer crawlers do not run JavaScript, so a client-rendered page can look blank to them. Check your AI crawl logs and status codes before rewriting anything.

See your AI visibility on your own brand

Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.

Start free