All posts
AI visibility·1 August 2026·9 min read

Is Your AI Visibility Data Real? A Reproducible Test

Every AI visibility tool — including CrunchJunkie — measures answer engines through an intermediary: an automated browser session, or an API call. Both can be served personalised or cached responses, and the two distort your numbers in opposite directions. Personalisation inflates variance; caching suppresses it, producing stable-looking charts that measure nothing. You can test for both in an afternoon, with the protocol below and no vendor cooperation required. We are a vendor writing about how to audit vendors, so read the “What we cannot prove” section before you cite anything here.

By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.
A CrunchJunkie AI-visibility score card reading 26.3% (plus or minus 4.8%), measured over 457 of 1,740 runs — a percentage shown with its sample size and margin of error.
A real AI-visibility number carries its receipts: 26.3% ± 4.8%, measured over 457 of 1,740 runs. A bare percentage with no sample size and no margin of error is asking for trust it hasn't earned.

Why measurement architecture is the whole ballgame

An AI visibility tool answers one question: when a buyer asks an answer engine about your category, how often does your brand appear? To answer it, the tool has to ask the engine — and there are only two ways to do that. The first is UI capture, or browser automation: the tool drives a real browser session against the consumer product at chatgpt.com or perplexity.ai. Peec AI, a Berlin-based competitor founded in early 2025, is one public example — it describes measuring answer engines by simulating browser sessions rather than by calling APIs. The second is API sampling: the tool calls the provider's API endpoint and records the response. Why the distinction is not pedantry: the consumer product and the API are different systems. The consumer app carries its own system prompt, memory, personalisation, model routing and retrieval pipeline. The API carries none of that. Perplexity's own documentation and third-party testing both note that the Sonar API behaves differently from the consumer product, and Google's AI Mode and AI Overviews have no public API at all — so browser capture is the only option that exists for them. Because brand visibility in a generative answer is overwhelmingly determined by what the engine retrieved before answering, and retrieval differs between surfaces, the surface you measure is not an implementation detail. It is the measurement. The honest conclusion is not that one architecture wins — it is that each is exposed to a different set of distortions, and a buyer should know which. We have argued the broader case that sampling design, not capture method, decides accuracy; this piece is the hands-on test.

The two failure modes

Each architecture is exposed to a different set of distortions. Here are the six that matter, and which capture method each one hits — the rows marked “Both” reach API sampling too, which is where the dangerous one lives.
Failure modeWhat it does to your dataSymptom in the dashboardPrimary exposure
Account memory / personalisationConditions answers on the measuring account's prior behaviourUnexplained jumps; two tools disagreeUI capture
Locale routing via proxyChanges the underlying web results, sometimes language and modelTrend shifts with no content changeBoth
Experiment bucketing (A/B tests)Different accounts land in different product variantsCross-account divergence read as brand movementUI capture
Answer-level cachingReturns a stored response to a repeated promptFlat line, near-identical textBoth
Retrieval-layer cachingFresh generation over stale retrieved documentsYour new content doesn't register for daysBoth
Bot mitigation / throttlingDegraded or truncated answers instead of a clean errorQuiet quality decay, no alertUI capture

Why one failure mode is far more dangerous

Personalisation is the noisier problem, and the safer one, because noise is visible. Someone eventually asks why share of voice jumped eleven points in a week. Caching is the dangerous one. A cache hit produces a flat, stable, professional-looking trendline — and in this market flatness is misread as data quality. A client reviews steady visibility and concludes their GEO programme is not working, when in fact nothing was measured. That is the bug that survives longest, because it never looks like a bug. There is a specific, testable consequence worth stating plainly: if answers are non-deterministic and yours are not varying, you are reading a cache. Generative models sample from a distribution and retrieval shifts hour to hour, so genuine repeated measurement of the same prompt should produce different phrasing and, often, a different set of named brands. Stability at the level of individual responses is not precision. It is a symptom.

The experiment: why the control brand is Semrush

Three tests follow. All three run against the consumer interface, which is what most tools claim to represent — run them yourself before you trust any vendor's chart, ours included. Every test uses Semrush as a control brand, and that choice does the methodological work. Semrush is one of the most established names in SEO tooling, with a large volume of third-party coverage, review-site presence and comparison content, so for broad category prompts it should appear at or near ceiling visibility. That gives you a known expected value. This matters because it separates two things that otherwise look identical. If your control brand's measured visibility swings from 100% to 60% across runs of the same prompt in the same hour, that swing is measurement noise — Semrush's actual market position did not move in forty minutes. Any tool reporting your own brand's movement at a magnitude smaller than that observed noise floor is reporting nothing. Substitute the dominant incumbent in your own category if you are not testing SEO tools; the only requirement is that the brand be near-certain to appear.

Test 1 — the repeat test (detects answer-level caching)

Setup. Open ChatGPT and Perplexity, disable memory, and use a temporary or incognito chat. Open a brand-new conversation for every single run. This is not optional: reuse a thread and run three is conditioned on runs one and two, and you will measure your own conversation instead of the engine. Procedure. Send the prompt below 20 times across roughly 10 minutes, in 20 fresh conversations. For each run, record: (a) the full list of brands named, (b) whether Semrush appears, (c) Semrush's position in the list, and (d) the first fifteen words of the response verbatim. Reading the result: • Opening sentences near-identical across most runs → answer-level caching, or something functionally equivalent. Your data has a resolution problem. • Brand list identical every run but phrasing varies → generation is fresh, retrieval is cached. This is the more common pattern and the more consequential one, because it is retrieval that determines who gets named. • Both phrasing and brand list vary → uncached. Now count how much they vary; that variance is your noise floor, and it is the number the rest of your reporting has to clear. Write the noise floor down as a percentage. If Semrush appeared in 17 of 20 runs, your single-measurement error bar is roughly ±15 percentage points — and a vendor reporting that your visibility moved from 22% to 26% on daily sampling is inside that band. The exact prompt, sent byte-identical each time:
Prompt
What are the best SEO tools for a mid-sized marketing agency in 2026?

Test 2 — the canary test (detects retrieval staleness)

Answer-level caching is invisible if you only ask questions whose answers do not change, so construct a question whose answer you control and whose timing you know exactly. Setup. Publish a page on your own domain containing a specific, unusual, verifiable factual claim — a dated statistic, a named methodology, a number that appears nowhere else. Record the publication timestamp, submit the URL via IndexNow, and confirm it is indexed. Reading the result. The lag between confirmed indexing and the engine reproducing the claim is your retrieval TTL for that engine, and that number is the true latency of your entire GEO feedback loop. If it is eleven days, then any test you run on new content inside eleven days returns a false negative, and any vendor reporting daily granularity on content changes is reporting inside their own lag. Run the canary separately per engine — they do not share infrastructure and, in our experience, they do not share latency. Two prompts do the work here:
PromptRun daily from publication:
According to [your domain], what is [the specific claim on the canary page]?
PromptThen, once it resolves, the harder version:
What sources report [the specific claim]?

Test 3 — the divergence test (detects personalisation and locale effects)

Setup. You need four conditions run simultaneously on the same prompt: 1. Account A, memory on, having previously run 30+ prompts about SEO tooling 2. Account B, memory off, no prior category history, temporary chat 3. Account B via a US IP 4. Account B via a German IP Run each condition 10 times, recording Semrush's appearance rate and the full brand list per condition. Reading the result: • A diverges from B → account memory is contaminating measurement. This is the industrialised version of a mistake individuals make constantly: searching your own brand repeatedly, then reading the personalised result as market data. At scale, a scraping account that has run thousands of category prompts is no longer a neutral observer — it has a profile. • US diverges from DE → locale routing is a live variable, and anyone reporting a single global visibility number is averaging over it silently. If you sell in multiple markets, insist on per-market reporting. • Free-tier diverges from paid → tier is a variable too, and neither tier is “what users see,” because real users are split across both. The prompt, identical across all four conditions:
PromptIdentical across all four conditions:
Which SEO and marketing reporting tools would you recommend, and why?

Seven questions to ask any vendor

These are answerable, and they discriminate. A vendor that answers all seven has thought about measurement; most cannot answer four. 1. Is account memory disabled, and is every prompt run in a fresh conversation? 2. Are measurement accounts rotated, and how is category contamination prevented? 3. Which subscription tier is measured, per engine? 4. Residential or datacentre IPs, and which locales? 5. Are prompts sent byte-identical, or paraphrased to defeat caching? 6. How many runs per prompt, per period? Not “daily” — the number of runs. 7. Is variance or a confidence interval reported alongside the mean? Question six is the one the industry avoids. “Daily tracking” is a cadence, not a sample size. One run per prompt per day, given the noise floor you measured in Test 1, cannot support a claim about a four-point weekly change; fifty runs can. Browser automation costs meaningfully more per query than an API call, which creates real pressure toward lower sample counts — so for any UI-capture tool this is the question that matters most. It applies to us too: we avoid running our own browser fleet — for the engines with no API we retrieve the rendered answer through a third-party search API, and everywhere else we use the providers' own APIs — but the runs-per-prompt question still lands on us, and you should hold us to it. Question seven is the tell. A tool that reports a mean without a confidence interval, on a non-deterministic system, is presenting a point estimate as a fact.

What we cannot prove

We are a vendor, so here is the boundary of what is verifiable in this article. Provable and publicly documented: that consumer apps and APIs are different systems with different retrieval; that ChatGPT offers memory and chat-history referencing that condition responses; that Perplexity's Sonar API behaves differently from its consumer product; that Google AI Mode and AI Overviews have no public API; that generative outputs are non-deterministic; and that Peec AI is Berlin-based, founded in early 2025, and publicly describes using browser-session capture rather than API calls. Not publicly documented, and therefore framed here as testable rather than asserted: whether any specific provider caches answer-level responses, what any provider's retrieval TTL is, and what any specific vendor's per-prompt sample rate is. We do not know Peec AI's internal sampling design, and we have not audited their implementation. Nothing above is a claim that their data is cached or personalised — it is a claim that the architecture is exposed to those effects, as is ours, and that the exposure is measurable from outside. Where CrunchJunkie is exposed. API sampling is immune to account memory and to UI experiment bucketing. It is not immune to locale routing, and it is not immune to retrieval-layer caching, which sits below both architectures. For the engines with no API — Google AI Overviews, Google AI Mode and Microsoft Copilot — we retrieve the rendered answer through a third-party search API rather than running our own logged-in browser sessions, which keeps account memory and experiment bucketing out of even those results, but leaves locale routing, retrieval-layer caching and a dependency on that provider's own capture. Any tool claiming architectural immunity here is overclaiming — and you can catch it with Test 1 in twenty minutes. If you run these tests against CrunchJunkie and the numbers do not hold up, send us the data; that is a more useful conversation than a comparison page.

Frequently asked questions

No. These systems retrieve documents and synthesise an answer; query volume for a brand is not a known input to that retrieval decision, and inference-time queries do not update model weights. Self-searching does have one reliable effect: it personalises your own account, which corrupts your ability to see what a cold prospect sees.

For fidelity to the consumer experience, yes — and for engines with no API it is the only option. But “exactly what users see” is overstated: a memoryless automated session on a datacentre IP is also an artificial condition, just a different one. It is higher fidelity, not ground truth, and sampling design usually affects data quality more than capture method does.

Because the underlying system is non-deterministic and its retrieved sources change continuously. Real measurement of a real system produces variance. Excessive smoothness suggests caching, insufficient sampling, or aggregation hiding both.

Enough that your reported changes exceed the noise floor you measured in Test 1 — a threshold that is category-dependent, which is why the test comes before the number. If your noise floor is ±15 points on single runs, one daily run cannot support a claim about a four-point move.

Automated access to consumer interfaces generally runs against provider terms. The practical consequence is usually operational rather than legal: bot detection, rate limiting, and silent data gaps when an interface is redesigned. Ask any vendor how those gaps are disclosed in reporting.

See your AI visibility on your own brand

Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.

Start free