All posts
AI visibility·23 September 2026·10 min read

We ran the same AI-visibility study twice. Eight of 19 agencies changed rank, and none of them moved.

On 8 September we published a ranking of US GEO agencies built from 198 AI answers. The obvious question about a ranking like that is whether it holds still. So we ran the identical measurement again, over a later stretch of days that shares no dates with the first, and put the two side by side. Eight of the nineteen agencies changed position. First and second swapped. And not one of the nineteen moved by an amount that samples of 168 and 90 answers can separate from chance. If you report AI visibility to clients, the distance between what the chart shows and what the data supports is the whole problem.

By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.

The short version

Two windows, no shared days: 4 to 8 September and 14 to 23 September. Same six buyer questions, same two engines, same nineteen agencies. Each agency was named in some share of 168 answers in the first window and some share of 90 in the second. Eight agencies changed rank between them. iPullRank went from first to second having lost 0.2 percentage points. Siege Media took first place on a gain of 4.4. First Page Sage climbed from ninth to sixth. Yes Optimist slipped from thirteenth to fourteenth. None of those movements is distinguishable from sampling noise. Every agency's 95% range in the first window overlaps its range in the second — all nineteen of them. The largest raw change in the table, Directive's 6.5 points, sits comfortably inside what a sample this size produces when nothing has happened at all.

The same nineteen agencies, measured twice

Read this as two independent estimates of the same quantity, not as a before and after. The ranges are 95% Wilson intervals: the band inside which the real rate plausibly sits, given how many answers stand behind it. Where one agency's two bands overlap, the difference between its two windows is not established. The second window is smaller, 90 answers against 168, so its bands are wider. That is not a defect in the second window. It is what fewer answers buys.
Share of AI answers naming each agency, measured over two windows with no shared days. ChatGPT and Claude pooled, six buyer questions, United States.
Agency4–8 Sep (168 answers)95% range14–23 Sep (90 answers)95% rangeChange
iPullRank36.9%30.0–44.436.7%27.5–47.0−0.2 pp
Siege Media34.5%27.8–42.038.9%29.5–49.2+4.4 pp
Omniscient Digital30.4%23.9–37.730.0%21.5–40.1−0.4 pp
Directive28.0%21.7–35.234.4%25.4–44.7+6.5 pp
Obility24.4%18.5–31.421.1%14.0–30.6−3.3 pp
19.6%14.3–26.317.8%11.2–26.9−1.9 pp
Go Fish Digital19.1%13.8–25.717.8%11.2–26.9−1.3 pp
NoGood19.1%13.8–25.718.9%12.1–28.2−0.2 pp
First Page Sage16.1%11.3–22.421.1%14.0–30.6+5.0 pp
Seer Interactive14.9%10.3–21.017.8%11.2–26.9+2.9 pp
WebFX14.9%10.3–21.014.4%8.6–23.2−0.4 pp
Foundation Marketing10.7%6.9–16.311.1%6.2–19.3+0.4 pp
Yes Optimist8.9%5.5–14.24.4%1.7–10.9−4.5 pp
Digital Elevator7.7%4.6–12.86.7%3.1–13.8−1.1 pp
Bay Leaf Digital4.2%2.0–8.33.3%1.1–9.3−0.8 pp
Citant.ai3.0%1.3–6.82.2%0.6–7.7−0.8 pp
Grow and Convert1.2%0.3–4.21.1%0.2–6.0−0.1 pp
NP Digital1.2%0.3–4.21.1%0.2–6.0−0.1 pp
Are You Visible0.0%0.0–2.20.0%0.0–4.10.0 pp
Share of AI answers naming each agency, measured over two windows with no shared days. ChatGPT and Claude pooled, six buyer questions, United States.

Eight changed rank, and rank is the thing people read

Nobody opens a client report and studies a confidence interval. They look at the order, and the order is the least stable thing on the page, because turning a rate into a position throws the margin away. Look at the top of the table. In the first window iPullRank led on 62 of 168 answers and Siege Media followed on 58. In the second, Siege Media led on 35 of 90 and iPullRank followed on 33. Four answers separated them the first time, two the second. One answer falling the other way on one day would have flipped it back. That is not a story about Siege Media taking the lead from iPullRank. It is a two-answer margin being reported as a change in market leadership. The same applies further down: First Page Sage's three-place climb rests on 27 of 168 becoming 19 of 90, two estimates whose ranges cover each other almost entirely. The agencies that genuinely separate are the ones far apart to begin with. iPullRank against Foundation Marketing is a real gap in both windows. iPullRank against Siege Media is not a gap in either.

How big a move has to be before it means anything

The useful form of this finding is a number you can apply to your own reporting. For a brand sitting near 30% visibility, here is how many answers each window needs before a change of a given size can be told from chance, at the conventional 95% confidence and 80% power. The first row is the one that matters. Detecting a five-point move takes roughly 1,377 answers in each window. At the rate this project runs — seven questions, twice a day, on two engines, so 28 answers a day — that is about seven weeks per window, fourteen weeks to compare two. Nobody in this category reports on that cadence. We do not either. The practical reading is not that you need fourteen weeks. It is that the movement in a weekly AI-visibility report is almost entirely noise, and that reporting it as performance is a claim the data cannot carry.
Answers needed in each window to detect a change from a 30% baseline, two-proportion test, 95% confidence, 80% power. Days assume 28 answers a day.
To detect a move of……each window needs…which is about
5 points, 30% to 35%1,377 answers7 weeks
10 points, 30% to 40%356 answers13 days
15 points, 30% to 45%163 answers6 days
20 points, 30% to 50%93 answers3 days
Answers needed in each window to detect a change from a 30% baseline, two-proportion test, 95% confidence, 80% power. Days assume 28 answers a day.

The engines disagree far more than the weeks do

Across the combined 4 to 23 September window each engine produced 129 answers per agency. Splitting the same answers by engine produces differences that dwarf anything the calendar produced. Seer Interactive was named in 5 of Claude's 129 answers and 36 of ChatGPT's. NoGood, 11 against 38. Omniscient Digital runs the other way, 50 against 28. Those gaps are wide enough to survive the sample size, which none of the week-to-week movements were. So when an AI-visibility number moves and you want to know why, the engine mix is a better first suspect than the market. Adding or dropping an engine, or a provider swapping the model behind a product name, shifts a blended figure further than any realistic change in how often people are actually recommended. We have written separately about why the engines disagree about your brand; this is the same effect, sized against the noise floor.
Answers naming each agency, by engine, 4–23 September 2026. Each engine produced 129 answers per agency. The eight widest gaps.
AgencyClaude (of 129)ChatGPT (of 129)Gap
Seer Interactive5 (3.9%)36 (27.9%)24.0 pp toward ChatGPT
NoGood11 (8.5%)38 (29.5%)21.0 pp toward ChatGPT
iPullRank35 (27.1%)60 (46.5%)19.4 pp toward ChatGPT
Omniscient Digital50 (38.8%)28 (21.7%)17.1 pp toward Claude
Go Fish Digital18 (14.0%)30 (23.3%)9.3 pp toward ChatGPT
Yes Optimist15 (11.6%)4 (3.1%)8.5 pp toward Claude
First Page Sage27 (20.9%)19 (14.7%)6.2 pp toward Claude
Bay Leaf Digital9 (7.0%)1 (0.8%)6.2 pp toward Claude
Answers naming each agency, by engine, 4–23 September 2026. Each engine produced 129 answers per agency. The eight widest gaps.

What we would put in a client report

Three changes, none of which needs a different tool. Report the count next to the rate. "Named in 33 of 90 answers" invites the right question in a way that "36.7%" does not. A client who can see the denominator can work out for themselves that a two-answer swing is not a trend. Put the uncertainty on the chart. A bar with a range on it is harder to over-read than a bar without one, and when two competitors' ranges overlap the honest line is that they are level, not that one is ahead. Stop reporting rank movement week to week. Report the rate, the count and the range, and let the order be whatever it is. If a client asks why the order changed, the answer is usually that it did not change in any sense worth acting on. None of this makes the number less useful. It makes the claims attached to it survive being checked, which over a retainer is worth more.

What this does not say

It does not say AI visibility cannot be measured. The engine differences above are real and large at this sample size, and so is the distance between an agency named in a third of answers and one named in none. Coarse facts come through clearly on a few hundred answers. Fine ones do not. It does not say the September ranking was wrong. Both windows put broadly the same five agencies at the top and the same seven at the bottom. The structure held. Only the order inside it moved. It does not generalise past this setup. Nineteen agencies, six questions, two engines, one market, twenty days. A category with more candidates, or with coverage that genuinely moves, could behave differently. The arithmetic about sample sizes is not specific to agencies, though. It applies to any brand whose visibility is measured as a share of answers, which is all of them.

Frequently asked questions

At the sample sizes most AI-visibility reports are built on, a few hundred answers or fewer, roughly 13 to 20 percentage points. For a brand near 30% visibility, telling a 5-point move from chance at 95% confidence and 80% power takes about 1,377 answers in each of the two windows being compared. On 198 answers the smallest detectable move is about 13.5 points; on 90 answers it is about 20.

Because rank is an ordinal view of a continuous number, and it discards the margin. Two agencies separated by two answers out of ninety will swap places routinely without either of them gaining or losing anything. In this study eight of nineteen agencies changed rank between two windows in which no agency changed by a distinguishable amount.

Only if it increases the number of answers inside the window you compare. Scanning daily and then comparing one day to the next leaves you with a smaller sample and a noisier number, not a better one. What helps is more questions, more runs per question, and comparing longer windows rather than adjacent short ones.

Report the level weekly if the client wants to see it. Do not report the week-on-week change as performance, because at a weekly sample size almost none of that change is real. A monthly or quarterly comparison over a pooled window carries far more of the information and far less of the noise.

Both, for their own users. Over 129 answers each in this study, Seer Interactive was named by ChatGPT in 36 and by Claude in 5, while Omniscient Digital was named by Claude in 50 and by ChatGPT in 28. Neither is an error. A blended score hides that split, so it is worth keeping the per-engine numbers visible alongside it.

Keep reading

See your AI visibility on your own brand

Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.

Start 14-day trial