Does a GEO score predict AI visibility? We tested ours on 81 agencies. It does not.
We build a GEO audit that scores a website from 0 to 100. A score invites an obvious question: do websites with a higher score get named more often in AI answers? We had the data to check. For 81 German SEO agencies we knew how often four AI engines named them, and we could run the audit on every one of their websites. The answer is no, or as close to no as 81 websites can show. This post has the numbers, the limits of the study, and the full data as a file.
By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.
Each dot is one agency. Across: the GEO score of its website on 27 September 2026. Up: how many of 633 AI answers named it, 12–19 September 2026. Rank correlation +0.10.
The short version
The 34 agencies that were never named in 633 AI answers have an average GEO score of 84.3. The 47 that were named at least once average 86.6. The 19 that were named in ten or more answers average 84.3 again. The rank correlation between score and mentions is +0.10 on a scale from −1 to +1, and with 81 sites the plausible range runs from −0.12 to +0.32, which includes zero.
Of the five parts of the score, one leans positive: content, at +0.20. Crawler access, structured data, technical hygiene and llms.txt are all within 0.05 of zero. So the score does not tell you who gets named. What it does tell you comes further down. It is less exciting and more useful.
What we measured
Two measurements, taken independently of each other.
The first is visibility. Ten German-language questions a buyer would ask were put to ChatGPT, Gemini, Perplexity and Google AI Overviews, twice a day per engine, from 12 to 19 September 2026. That gave 633 answers. For each agency we counted the answers that named it in the answer text. It is the measurement behind our ranking of German SEO agencies in AI answers, in its first window.
The second is the audit. On 27 September 2026 we ran the GEO audit on each agency's website, with the same code the product runs. It reads robots.txt, requests the homepage with the user agents of the AI crawlers, and samples up to ten pages, taken from the sitemap where there is one. 78 of the 81 sites were audited on ten pages. The score is a weighted sum of five parts: AI crawler access (30 points), content accessibility (30), structured data (20), technical hygiene (15) and llms.txt (5).
We started with 87 agencies that have a website on record in the project. Six were left out because their homepage did not answer our request normally that evening: a 503, a 403, a bot challenge, no connection. A failed request is not a measurement of a website, so those six are not in any number here. That leaves 81.
Sorted by mentions: the score barely moves
Start from the outcome. If the score predicted visibility, the agencies the engines name should score clearly higher than the ones they never name. They do not. Between the never-named and the named there are 2.3 points, and the 19 agencies with ten or more mentions fall back to exactly the average of the never-named.
The ranges overlap almost completely. Never-named agencies score between 67 and 96. The eight agencies named in fifty or more answers score between 71 and 96. Among the agencies that were named at all, more mentions even go with a slightly lower score (correlation −0.18, 47 agencies), which is as meaningless as the slightly positive figure for the whole set.
GEO score of 81 German SEO agencies, grouped by how many of 633 AI answers named them (12–19 September 2026). The last two groups are part of the second.
Group
Agencies
Average score
Median
Lowest to highest
Never named
34
84.3
85
67–96
Named at least once
47
86.6
88
65–98
Named in 10 or more answers
19
84.3
87
65–96
Named in 50 or more answers
8
86.4
89
71–96
GEO score of 81 German SEO agencies, grouped by how many of 633 AI answers named them (12–19 September 2026). The last two groups are part of the second.
The ten most-named agencies and their scores
The ten agencies the engines named most often, with the score of their website. Scores from 71 to 96 sit side by side at the top of the visibility ranking. A score is a statement about a website on one evening, not about the quality of an agency's work.
The ten agencies named most often in 633 AI answers, 12–19 September 2026, with the GEO score of their website on 27 September 2026.
Agency
Answers naming it
Share of answers
GEO score
Claneo
200
31.6%
88
Aufgesang
178
28.1%
87
SEO-Küche
115
18.2%
76
Bloofusion
105
16.6%
71
Online Solutions Group
97
15.3%
91
morefire
91
14.4%
92
eology
76
12.0%
96
xpose360
74
11.7%
90
DREIKON
43
6.8%
77
HECHT INS GEFECHT
37
5.8%
73
The ten agencies named most often in 633 AI answers, 12–19 September 2026, with the GEO score of their website on 27 September 2026.
Sorted by score: a lean, not a relationship
Now from the other side. Sort the 81 agencies by score and cut them into three groups of 27. In the lowest group 13 agencies were named at least once, in the middle group 15, in the highest 19. That is the lean behind the +0.10: a higher score goes with a slightly better chance of being named at all.
It does not go with being named more. The middle group collects the largest share of answers, 3.5%, and the highest group the smallest, 2.1%. Tested against chance, the correlation of +0.10 has a p-value of 0.37: shuffle the scores randomly across the agencies and you get a figure at least this large more often than one time in three.
The same 81 agencies sorted by GEO score and cut into three groups of 27. Share of answers is pooled: all mentions of the group divided by all answers (633 × 27).
Score
Agencies
Named at least once
Mentions in total
Share of answers
65–83
27
13
372
2.2%
84–90
27
15
601
3.5%
91–98
27
19
355
2.1%
The same 81 agencies sorted by GEO score and cut into three groups of 27. Share of answers is pooled: all mentions of the group divided by all answers (633 × 27).
The five parts of the score
The score is built from five parts, so one of them could carry a signal that the sum hides. Four do not. For crawler access there is a plain reason: 72 of 81 sites have full marks and the other nine all lose the same points for the same reason, the refused crawler request described below. A part that takes only two values cannot explain much.
Content is the exception, at +0.20 with a range from +0.01 to +0.40 and a p-value of 0.07. It is the only part where pages with more text, clearer structure and more sourced material go with more mentions, and it fits what one would expect. We would still not put weight on it. We looked at five parts, and one of five looking mildly positive is what chance produces fairly often. It is a reason to test content again on a second category, not a finding.
Rank correlation (Spearman) between each part of the GEO score and the number of answers naming the agency, 81 agencies. The range is the middle 95% of 2,000 resamples.
Part of the score
Points of 100
Average result
Correlation with mentions
Plausible range
AI crawler access
30
98%
+0.05
−0.18 to +0.24
Content accessibility
30
86%
+0.20
+0.01 to +0.40
Structured data
20
73%
0.00
−0.22 to +0.22
Technical hygiene
15
92%
+0.05
−0.17 to +0.26
llms.txt
5
47%
+0.05
−0.18 to +0.27
Whole score
100
85.7
+0.10
−0.12 to +0.32
Rank correlation (Spearman) between each part of the GEO score and the number of answers naming the agency, 81 agencies. The range is the middle 95% of 2,000 resamples.
Why there is no check-by-check league table
The obvious next step is to compare, for every single check, the visibility of sites that pass with the sites that do not, and to publish the checks with the biggest difference as the ones that matter. We ran that comparison and decided against presenting it as a result. Here is why.
Mentions are concentrated. Eight agencies account for 936 of the 1,328 mentions in the set, which is 70%. Any split of 81 sites into two groups is therefore mostly a statement about which side those eight landed on. Take the title tag: the 17 sites whose titles are between 10 and 60 characters on every sampled page were named in 0.4% of answers, the other 64 in 3.2%. Nobody should conclude that long titles get an agency recommended. The same goes for llms.txt, where the 40 sites that publish the file were named in 1.7% of answers and the 41 without it in 3.4%.
One example from our own work. The first run of this study audited mostly homepages, because of a bug in how the audit read sitemap indexes. In that run, sites with sourced statistics were named more often, and it looked like a finding. After the bug was fixed and the audit read ten pages per site, only five sites passed that check on every page, and they were named less than the rest (0.4%). The finding was gone. Every column is in the file below, so you can run the comparison yourself. Treat what comes out as a question for a larger sample.
What the audit found on agency websites
The audit results are worth a look on their own, because these are the websites of people who optimise websites for a living. Overall they are in good shape: 62 of 81 score 80 or more, none scores below 65.
All 81 allow the AI search crawlers in robots.txt. Nine of them nevertheless answered a request carrying an AI crawler's user agent with a refusal while the same page loaded normally for a browser. A caveat on that number: our requests came from an ordinary address, not from the crawler operators' networks, so a firewall that checks where a crawler comes from would refuse us and let the real one in. Nine is an upper limit. One of the nine is among the three most-named agencies in the set, which says something about how little the agency's own site has to do with being named.
40 of 81 publish an llms.txt, 41 do not. 17 have no Organization schema at all and only 27 pass the check for a complete one with name, url, logo and sameAs. 24 have no breadcrumb schema. On five sites some of the sampled pages only show their content after JavaScript has run. The content checks taken from research are rarely met across a whole site: 10 sites have direct quotations on every sampled page, five have sourced statistics on every page, and on 44 sites not a single sampled page links to an authoritative source.
What the audit found on the websites of 81 German SEO agencies, 27 September 2026, up to ten pages per site. For checks made on every page, a site passes when all sampled pages pass and fails when none does.
Check
Pass
Partly
Fail
AI search crawlers allowed in robots.txt
81
0
0
AI crawlers get through the firewall
72
0
9
Content is in the HTML without JavaScript
76
5
0
At least 250 words of text
62
19
0
One H1, headings nested in order
8
73
0
Lists or data tables
69
12
0
Direct quotations
10
69
2
Statistics with a source
5
74
2
Links to authoritative sources
0
37
44
Valid JSON-LD
67
11
3
Organization schema with name, url, logo, sameAs
27
37
17
Breadcrumb schema
47
9
24
Self-referencing canonical
66
8
7
Title tag of 10 to 60 characters
17
64
0
Meta description of 50 to 160 characters
36
43
2
llms.txt present and valid
39
1
41
What the audit found on the websites of 81 German SEO agencies, 27 September 2026, up to ten pages per site. For checks made on every page, a site passes when all sampled pages pass and fails when none does.
Why a good score does not get you named
Because the answer to a category question is mostly decided somewhere else. When we counted the sources the engines cited for this same category over ninety days, one best-agencies list on a third-party site was cited in 345 of 1,033 answers, and rankings, directories and review pages filled most of the top ten. That analysis is in the pages AI already cites. An engine that is asked which SEO agency to hire reads lists and names who is on them. How well the agency's own site is built has little to do with it.
This is the difference between being cited and being named, which we have written about in cited is not named. An audit looks at whether a site can be fetched, read and used as a source. Whether a brand is named depends on what the pages the engine trusts say about it.
It also explains how this result sits next to the research the content checks are based on. The GEO paper by Aggarwal and colleagues found that adding statistics, quotations and sources to a page raised that page's visibility inside generated answers in their experiments. Their question was whether a page gets used more once it is among the sources. Ours is whether a brand gets named. The two can both be true. This study does not contradict theirs, and it does not confirm it for brand mentions either.
What a GEO score is for
A floor. The audit finds the things that can keep a site out without anybody noticing: a firewall rule that refuses AI crawlers, content that only exists after JavaScript has run, no schema that says who the organisation is, a canonical pointing somewhere else. None of these gets a brand named. Each of them can stop a page from being read or used. They are cheap to check and most are cheap to fix.
So the order of work we would suggest is this. Run the free GEO audit and fix what blocks. Do not spend a month moving a score from 85 to 95, because nothing in this data says the engines will notice. Then measure the thing you care about directly: ask the engines your buyers' questions, repeatedly, and count who gets named. The free AI visibility check does a first pass of that. And then work on where the engines get their names from, which is mostly other people's pages.
What this does not say
One category, one country, one week of answers, one audit run. German SEO agencies are a population whose own websites are well maintained. Scores run from 65 to 98 and three quarters of the sites are at 80 or above. The study says nothing about a site that scores 40. In a population with broken sites the score may well separate the visible from the invisible, because a site no crawler can read has a different problem from the ones measured here.
The outcome is brand mentions in answers to category questions. We did not test whether the score predicts how often an agency's own pages are cited as a source, which is closer to what an audit looks at. That is the next thing to measure.
Visibility is very unevenly spread: 34 of 81 agencies have no mention at all and eight have 70% of them. With that shape, 81 sites cannot detect a weak relationship. The range of −0.12 to +0.32 means a modest positive effect is still possible. A strong one is not.
The comparison is observational. Nothing was changed on any site to see what happens next. And we sell the audit and the measurement, so we have an interest in both. That is the reason the data is published with the post.
For the same reason, here is our own result. The website of Tiki-Taka Media, the agency Philipp Enders runs, scored 98 in the free version of the same audit on 29 September 2026, and the result is public. That is level with the highest score among the 81 agencies. Tiki-Taka is not one of the 81, because it is not on the list of agencies the answers were counted against, so there is no mention figure to set beside the score.
The data
The file is here: geo-score-vs-ai-visibility-2026-09.csv. It has one row per agency, 81 rows, under CC BY 4.0. Use it, recompute it, publish what you find, and name the source.
Each row has the agency and its domain, the number of answers that named the agency out of 633, the GEO score and its band, the number of pages audited and where they were taken from, the result for each of the five parts of the score from 0 to 1, the separate agent-readiness score, and the outcome of 23 single checks as pass, partial or fail.
To check a row, or to add your own site to the comparison, run the free audit on the domain. Scores change when sites change, and a firewall may treat a request differently from one day to the next, so a repeat of the audit lands near the first figure and not exactly on it.
Frequently asked questions
Not in the data we have. Across 81 German SEO agencies, the GEO score of the website and the number of AI answers naming the agency had a rank correlation of +0.10, with a plausible range from −0.12 to +0.32. Agencies that were never named in 633 answers scored 84.3 on average, agencies named at least once 86.6.
It finds blockers: AI crawlers refused by a firewall, content that needs JavaScript to appear, missing organisation schema, broken canonicals. These do not get a brand named, but they can stop a site from being read or used as a source. Fix them once, then measure visibility directly by asking the engines.
This data shows no relationship. 40 of 81 agencies publish an llms.txt and 41 do not, and the correlation between the llms.txt part of the score and mentions is +0.05. Of the agencies with the file, 27 of 40 were named at least once, against 20 of 41 without it. The agencies without it collected the larger share of answers, 3.4% against 1.7%. The two figures point in opposite directions, which is what no relationship looks like.
Yes. The file has one row per agency with the mentions out of 633 answers, the GEO score, the five category results and the outcome of 23 checks. It is published under CC BY 4.0 at https://crunchjunkie.io/blog/data/geo-score-vs-ai-visibility-2026-09.csv.