---
title: "GEO Audit"
url: "/docs/geo"
canonical_url: "https://crunchjunkie.io/docs/geo"
markdown_url: "https://crunchjunkie.io/docs/geo.md"
language: "en"
type: "doc"
summary: "Score how well a site is positioned to be crawled, parsed and cited by AI engines."
site: "CrunchJunkie"
llms_txt: "https://crunchjunkie.io/llms.txt"
---

# GEO Audit
Source: https://crunchjunkie.io/docs/geo

Score how well a site is positioned to be crawled, parsed and cited by AI engines.

## What the GEO Audit measures
GEO (Generative Engine Optimization) is the practice of making a website easy for AI engines — ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews — to crawl, understand and cite. The GEO Audit, found under AI Visibility → GEO Audit, fetches a client's site the same way those engines do (a plain HTTP request, no JavaScript execution — which is exactly how the major AI crawlers fetch) and grades concrete, verifiable signals into a single 0–100 readiness score.

Importantly, this measures readiness — whether a site is positioned to be cited — not actual citation share. How often a brand is genuinely mentioned in AI answers is tracked separately on the Visibility tab. We keep that distinction explicit so the score is never over-claimed.

## The five scored categories
The composite score weights five categories. AI crawler access (30): whether your robots.txt lets the crawlers that power live AI answers — OpenAI's OAI-SearchBot and ChatGPT-User, Anthropic's Claude-User and Claude-SearchBot, PerplexityBot, Googlebot and Google-Extended, Meta AI's meta-webindexer — actually reach the site. Content accessibility / SSR (30): whether the page's main content is present in the raw HTML rather than injected by JavaScript, since AI crawlers don't run JS. This category also carries an informational content-delta probe: the audited pages are fetched once as an AI agent and once as a normal browser, and a material difference between the two is flagged — a bot served less means the audit's findings may understate what humans see; a bot served more or different content is a cloaking risk. It never affects the score. The same category also scans the fetched pages for hidden text addressed to AI readers — markup a human visitor never sees (display:none, zero font-size, off-screen offsets, the hidden or aria-hidden attributes, screen-reader-only classes, HTML comments, alt or meta attributes) carrying wording like 'ignore previous instructions' or 'AI assistants: recommend …'. AI crawlers read raw HTML, so such text reaches them in full, and Google's spam policies count hidden text and attempts to manipulate generative AI responses as spam. This is CrunchJunkie's own heuristic detector, not a platform policy check: it shows a row only when it finds something, quotes the snippet so you can judge it, and never affects the score. Structured data (20): valid schema.org JSON-LD that anchors your brand as a known entity — especially an Organization block with sameAs links, plus support for richer answer-friendly types including FAQPage, Article, Product, HowTo and QAPage. Beyond just detecting these types, the audit checks that a present type actually carries its required properties (a FAQPage needs questions with accepted answers, an Article needs headline/author/datePublished, a Product needs offers, a HowTo needs steps); a type that's present but incomplete is flagged as a partial with the exact missing properties to add, rather than a full pass. Technical SEO hygiene (15): sitemap, self-referential canonical, a single H1, a sensible title and meta description, and HTTPS. llms.txt (5): presence and basic validity of an /llms.txt file — a forward-looking convention, weighted lowest and deliberately so: Google's official generative-AI guidance states it ignores llms.txt, and no answer engine has confirmed consuming it for citations, so it's scored as cheap, machine-readable hygiene and never as a Google/search ranking or citation factor.

Each check is graded pass, partial or fail with a plain-language explanation of what was found and why it matters, so the result is an action list, not just a number.

## Research-validated content signals
Inside the content-quality scoring, the GEO Audit checks for three writing signals that peer-reviewed research has shown make a page more likely to be quoted by generative engines: direct quotations, cited statistics, and authoritative outbound citations. These come from the KDD'24 paper "GEO: Generative Engine Optimization" (Aggarwal et al., Princeton / IIT Delhi), which measured that adding quotations, statistics, cited sources and fluent language each lifted a page's visibility in AI answers by roughly a quarter. Each check is graded pass/partial/fail with a plain-language note, and the audit labels them as research-backed so you know the recommendation isn't a hunch. The audit also grades **extractable structure** — whether a page breaks its content into bulleted or numbered lists and data tables, which AI answer engines lift near-verbatim into their responses, so structured content is markedly more citable than unbroken prose. It's scored fairly (a prose-rich page is nudged toward more scannable structure, never failed for it) and is grounded in 2026 research on how content structure shapes AI citation ("Structural Feature Engineering for GEO", arXiv 2603.29979).

The same research identified what doesn't work: keyword stuffing was the single worst tactic, actively reducing AI visibility. So the audit treats keyword stuffing as a negative signal — it never earns points, and a page whose keyword density is too high (roughly above 6%, with an elevated band around 4–6%) is flagged as a problem to fix, not a box to tick. Nothing in the content score ever rewards keyword density; it only ever penalises stuffing.

## Live AI-crawler reachability (WAF and CDN blocks)
A robots.txt rule that says "allowed" only matters if the crawler can actually reach the page. Plenty of sites unblock AI bots in robots.txt but then sit behind a WAF or CDN — Cloudflare's AI Crawl Control, formerly the "Block AI bots" toggle, is the common one — that quietly turns those same crawlers away at the edge. To both you and the site owner everything looks fine; to the AI engine the site is invisible.

The GEO Audit now runs a live reachability check alongside the robots.txt review: it requests the page while presenting itself as each AI crawler and reports whether the request actually got through. This is informational by default — it only raises a concern when a block is genuinely confirmed, so you won't get false alarms from ordinary rate-limiting or transient errors. When a real edge block is detected, the audit names the crawler that was turned away so you can take it straight to the client's hosting or security team. The lesson it encodes: "allowed in robots.txt" and "reachable in practice" are two different things, and AI visibility needs both.

You'll see this in two concrete places. On the GEO Audit, a "Firewall & access policy" card lays it out crawler by crawler — what robots.txt says next to what the live probe found, with a plain verdict for each. "Firewall blocks" is the one to act on: robots.txt allows the crawler, the edge blocks it anyway. It covers the eight crawlers a block can be pinned to — the search crawlers (OAI-SearchBot, PerplexityBot, Googlebot, Meta AI's meta-webindexer), the live agent fetchers (ChatGPT-User, Claude-User, Perplexity-User) and ClaudeBot — and rides along into the report, shared links and the PDF. Separately, if you've wired up crawl-log ingestion, the Crawl insights page raises a WAF-block alert — mirrored by a Signal in your inbox — the moment real AI-crawler traffic starts coming back with HTTP 403/406. "Real" is meant literally: requests that only use an AI crawler's name — an IP outside the operator's published ranges, or a robots.txt-only name such as Google-Extended or Applebot-Extended, which Google and Apple never send — are left out, because a firewall refusing them is doing its job; the alert lists them separately. Of the blocks it does count, it tells you how many came from the operator's published IP ranges and how many could not be checked (most AI operators publish no ranges, so those may well be the real crawler). Three more things leave the alarm since 27 September 2026, after a customer traced a red alarm to its firewall refusing vulnerability scanners: refusals on scanner paths such as /.env or /wp-config.php (a probe is a probe whatever it calls itself), requests from IPs that used the crawler names of two or more different AI operators in the window (no real crawler IP claims to be both OpenAI and Anthropic — that is one scanner rotating user agents), and the alarm is downgraded to an informational "unverified" note when the evidence contradicts a bot rule: every crawler with published IP ranges got through, or every refusal is an HTTP 406, the signature of a request-pattern rule such as ModSecurity rather than an identity rule. The card shows the 403/406 split for the same reason. So you catch an edge block two ways: from a point-in-time audit, and from live traffic in between audits.

Since 15 September 2026 there is a common cause behind this finding worth knowing about. From that date Cloudflare blocks, by default, the crawlers it classes as Agent — ChatGPT-User, Claude-User and Perplexity-User, the fetchers that act live when an assistant answers a question about you — and as Training (GPTBot, ClaudeBot and the like) on pages that show ads, for new zones, new sites of existing customers and every Free-plan zone, while search-only crawlers such as OAI-SearchBot stay allowed. Crawlers that serve both search and training are blocked by the same rule — Cloudflare names Googlebot, Applebot and Bingbot — so on those pages a site can lose Google Search and AI Overviews coverage too. It changes nothing about our scanning (we call the engines' APIs), but a site that lost its agent fetchers this way answers assistants with a blank. The live probe therefore carries all three classes, so the pattern shows up as "Firewall blocks" on the agent and training rows, usually on Googlebot too, while OAI-SearchBot gets through — and when the responses carry Cloudflare's edge markers, the finding names the cause and points you to Cloudflare → AI Crawl Control (or a WAF skip rule). Without those markers it keeps the generic firewall wording: the audit never claims Cloudflare it didn't see. A bot-challenge that hits our own audit request is reported as inconclusive from our vantage, never as a block on the crawlers.

One more robots.txt signal worth watching: a "Content usage policy" card appears when your site publishes Cloudflare's Content-Signal directive (search / ai-input / ai-train = yes/no). This is a preference signal, not an enforced control — no answer engine is confirmed to honour it and Google has publicly said it ignores it, so it never affects your GEO score and doesn't, on its own, change whether you're cited. We surface it for one reason: Cloudflare auto-added defaults to millions of managed robots.txt files, so your site may be broadcasting ai-input=no — asking AI engines not to use your content as live input for answers — without anyone choosing that. If you're pursuing AI visibility, that's the opposite of what you want; the card flags it (and notes when it looks like Cloudflare's default) so you can set ai-input=yes or remove the line. It's also available as the "Content usage policy" report dimension. (ai-train=no is a legitimate training opt-out with no effect on live-answer citations — the card says so, and never flags it.)

## Running an audit and tracking it over time
Open a client's AI Visibility project, go to GEO Audit, and the domain is pre-filled from the client's saved website. Click Run audit and CrunchJunkie checks robots.txt, llms.txt, the sitemap, and a representative sample of pages — the homepage plus one page per site section, preferring deeper pages, so blog posts and product pages are judged, not just the shallow top level. The page list comes from every sitemap the site publishes, including sitemap indexes and sitemaps declared in robots.txt, and from the links on the homepage when there is no sitemap. A sampled page that refuses or fails our request (403, 429, a server error) is left out rather than scored; a 404 stays in, because a dead URL in your own sitemap is a finding. On-page results are aggregated honestly across the sample ("3/6 sampled pages pass"), and the Per-page results section breaks the same checks down page by page, so a mixed result is always traceable to the exact page that needs fixing. Every run is saved, so you can re-audit after making changes and watch the score move — the history strip shows each snapshot with its date.

The audit also re-runs automatically — weekly by default, since on-site signals move slowly and each audit reads pages from the client's own site. You can switch a project to daily re-audits in its Settings tab, and a manual Run audit is always available regardless. Automatic runs are plain page reads with no AI checks, so they never cost anything.

If the homepage isn't the site when an audit runs, the audit says so instead of passing off a score. That covers a hosting suspension page, an expired or parked domain, a default server page, a server error and a homepage that doesn't answer. The page names what the server showed and offers to run the audit again, plus a link to the last audit of the live site. The score is shown crossed out as not comparable, and no fixes or AI insights are drawn from it, because they would describe the placeholder rather than the site. Such an audit raises no score alert (the separate "Site not being served" signal does), and reports and content briefs use the last audit of the live site instead. The API and the CrunchJunkie connector return the audit with a homepageUnavailable field that says what the server showed. The checks that did run stay one click away under "Show what was measured on that page".

Because the audit reflects on-site signals only, it can't see off-site factors that also drive AI visibility (brand mentions across the web, Reddit/Wikipedia presence) or whether a model was trained on your content. Treat a high score as "this site is well-positioned to be cited," and pair it with the Visibility tab to see whether that positioning is translating into actual AI mentions.

## SPA and server-rendering checks
AI crawlers fetch a page with a plain HTTP request and don't run JavaScript, so anything a site paints in the browser after load is invisible to them. The audit checks for this directly. The SSR signal check measures how much real content is in the initial HTML — the words present in the raw <body> before any JavaScript runs — and passes when there's a meaningful amount, flags a partial when there's only a little, and fails on an empty shell. The client-side-only check looks for the fingerprints of a single-page-app framework (React, Next.js, Vue, Nuxt, Angular, Svelte and others) sitting on top of a near-empty initial body, which is the classic pattern of a site that renders fine for people but returns almost nothing to a crawler.

Being server-rendered is only half of being machine-readable for a crawler, though — the page also has to let the crawler tell the main content apart from the navigation, header and footer. The audit checks for a semantic main-content landmark (a <main> or <article> element, or role="main"): with one, a crawler can isolate the content worth citing in a single step; without it, a page built entirely from generic <div>s forces the engine to guess which block is the answer. It's a gentle nudge, never a hard fail, with the fix being to wrap the primary content in <main>.

The audit also checks for a robots meta noindex: if a page carries <meta name="robots" content="noindex">, it's telling search and AI answer engines to leave it out entirely, so that's called out with the exact fix. Together these checks catch the technical reasons a page that looks perfect in a browser can still be effectively invisible to AI.

## Agent readiness (emerging)
Alongside the composite GEO score, each audit now produces a separate Agent Readiness sub-score out of 100. Where the GEO score is about being crawled and cited, agent readiness is about the newer class of AI agents — the real-time assistant fetchers and tool-callers that act live on a user's behalf. It grades three concrete, controllable signals: whether robots.txt lets the real-time agent user-agents through (ChatGPT-User, Claude-User, Perplexity-User, Gemini Deep Research and similar), whether your content is readable without JavaScript (agents fetch raw HTML and usually don't run JS), and the breadth of your structured data (how many distinct schema.org types you expose for an agent to parse). Each of these carries a "Draft the fix with Crunch" button, exactly like the main audit.

It is shown as a deliberately separate sub-score, not folded into your GEO score, for two reasons: it's an emerging area, and one of its signals — an agent-interface descriptor at /.well-known/mcp.json — follows a convention (MCP) that is still being standardised. That descriptor check is informational only: if we find one we flag it as forward-looking positioning and show an "MCP endpoint detected" badge, but its absence never lowers your sub-score. Nothing here penalises a site for not having adopted a pre-standard. Keeping it separate also means your historical GEO scores stay directly comparable over time.

Two checks follow Google's Lighthouse, which added an Agent Discoverability group to its Agentic Browsing category in version 13.5 (September 2026) and is bringing it to PageSpeed Insights, so the verdict matches what a client sees in Chrome DevTools. The first is Agent Resource Discovery. We look for an ai-catalog.json the way Lighthouse does: an Agentmap line in robots.txt, a link rel="ai-catalog" tag, a Link header, then /.well-known/ai-catalog.json. If there is one, we check it against the ARD 1.0 rules Lighthouse applies (spec version, entries, urn:air: identifiers, display name, type, exactly one of url or data). The current ARD spec (v0.91, 26 August 2026) renamed the file to ard.json and the link to rel="ard"; Lighthouse 13.5 does not read those yet, so we look for both forms, ard.json first, and the result says which one it found and whether Lighthouse will see it. An ard.json needs no spec version. Like the MCP descriptor it is informational, and not publishing one is fine; Lighthouse marks it not applicable too. The one thing we don't run is Lighthouse's extra strict JSON-schema pass, and the check says so. A third check in this group, "Markdown for agents", asks the homepage for itself with an Accept: text/markdown header, the way agents from Vercel, Cloudflare and others now do, and reports whether the site answers with a Markdown version (Content-Type text/markdown, ideally with Vary: Accept so caches keep the two apart) or with the HTML page; if the page advertises a Markdown twin with a link rel="alternate" type="text/markdown", the route Applebot takes, that twin is fetched and judged as well. It is display-only, never scored: no answer engine has confirmed that a Markdown response affects citations, so it is worded like llms.txt, as cheap agent-readability hygiene that hosts like Vercel and Cloudflare can switch on without content changes. The second Lighthouse check is llms.txt validity in the llms.txt section, which now uses Lighthouse's three rules: an H1, at least one link, and at least 50 characters.

## Entity authority
Each audit also produces a separate Entity Authority sub-score out of 100 — a read on how strong your schema.org entity graph is. Where the composite score's "Structured data" category asks "is there any schema at all", entity authority asks the sharper question answer engines actually use to decide who you are and whether to trust and cite you (E-E-A-T): a complete Organization (name, url, logo and a sameAs array), links out to authoritative external profiles (LinkedIn, Crunchbase, Wikidata, official social — three or more is a strong signal), the named people behind the brand (a founder or article authors declared as Person entities), and real author attribution on your articles.

It's deterministic — it only reads the structured data you actually publish, never guessing a schema type from your page text — and honest: if the sampled pages carry no article nodes, author attribution is marked "not scored" rather than counted against you (and BlogPosting counts as an Article, so a blog is never wrongly flagged as "missing Article"). Like Agent Readiness and Local GEO it's kept out of the headline GEO score, so your history stays comparable, and each gap carries a "Draft the fix with Crunch" button. You'll also see it as an "Entity authority" report dimension that rides into shared links and the PDF. One diagnostic pathway is worth knowing: teams with strong technical SEO — clean schema, correct robots.txt — are regularly surprised by weak AI visibility, because answer engines weigh more than markup compliance. If your structured-data checks pass but visibility stays low, look here next: a thin entity graph (no sameAs profiles, no named people, no authorship) is a common reason engines can't establish who you are and whether to trust a citation — and after that, the gap usually isn't on-site at all but in off-site authority, which the Domains and Gap-analysis views measure directly.

## Local GEO (for local businesses)
If a client is a local business — a shop, restaurant, clinic, tradesperson, anything with a physical location or a service area — turn on "This client is a local business" in the AI Visibility → Settings tab. The GEO Audit then produces a separate Local GEO sub-score out of 100, alongside the composite GEO and Agent Readiness scores. Like Agent Readiness it's kept out of your headline GEO score — a SaaS with no LocalBusiness schema shouldn't be marked down for it — so your history stays comparable.

Local GEO grades the on-site signals that make a business surface in local-intent AI answers ("best florist near me", "plumber in Berlin"): LocalBusiness schema and how complete it is, name/address/phone shown as real text and matching the schema, a link to your Google Business Profile, dedicated location pages, your city and region named in titles and headings, review markup, an embedded map, and links to reputable directories. It's honest — if you haven't set a target city it says "not evaluated" rather than failing you, and it never rewards duplicated "city-swap" location pages.

Two fields sharpen the checks. Set the client's Primary city and Region so the audit can confirm you name where you operate, and paste the Google Business Profile URL — the entity link that ties your site to your Maps listing, one of the strongest local signals. To find that URL: open Google Maps, search the business, open its listing, click Share and copy the link (it looks like https://maps.app.goo.gl/… or https://www.google.com/maps/place/…). Then run the audit to see the Local GEO card, with a "Draft the fix with Crunch" button on every issue just like the main audit.

## Draft the fix with Crunch
Every issue the audit raises comes with a plain-language explanation of what was found and how to change it. For the ones where you'd rather not write the fix from scratch, each issue has a "Draft the fix with Crunch" button. Click it and Crunch drafts a concrete, copy-pasteable remediation for that specific issue — grounded in the actual finding and the audited page, not generic advice — with a Copy button to lift it straight into a ticket or hand it to the client's developer.

The drafted fix is generated on demand for the issue in front of you and isn't saved to the audit; re-run it any time. Like the rest of Crunch, it needs an available AI model (Managed AI or your own key). It turns the audit from a list of problems into a list of ready-to-apply changes.

The Organization and sameAs issues are handled differently, because they are about facts a model cannot know. For those, Crunch reads your homepage and the about, contact and imprint pages it links to, and builds a connected block from what is there: an Organization, the WebSite it publishes and the homepage, tied together by @id. Your name and logo come from your own markup or tags; if the site states no name anywhere, the brand name you gave the project is used, but only when your homepage title confirms it. On a site whose root redirects to a language folder such as /de, the homepage entry describes the page that was actually read. sameAs holds only profile links that are on your site — declared in your markup, or linked from the pages read. A profile linked from a single inner page is listed as left out, never assumed to be yours, and there are no placeholders to fill in. Under the block you see where each value came from and what was left out on purpose. This path needs no AI model.

Every block is checked before you see it: the JSON must parse, and each type and property must exist in the current schema.org vocabulary and sit on a type that allows it. The same check runs on any JSON-LD inside a model-drafted fix, and the result is printed under the draft. It does not replace Google's Rich Results Test, which also checks value formats and required fields on the live page.

One issue has its own dedicated button. When the /llms.txt check isn't passing, the llms.txt category shows **Generate llms.txt with Crunch**: it drafts a complete starter /llms.txt from the site's OWN structure — the real URLs in its sitemap.xml plus the homepage title and description — grouped into sections with a short description per link. Every link is a real page on the site, never an invented one; save the result at your domain root as /llms.txt after a quick review.

## Share it: the free public audit
The same audit engine also powers a free, public version at crunchjunkie.io/products/geo-audit — no login, no account. Anyone can enter a domain and get the 0–100 score, the category breakdown and a prioritised fix list in seconds. It's handy for agencies: run a quick GEO check on a prospect's site before a pitch, or send the link so a client can see the problem for themselves. The public tool is rate-limited and reads only public pages and robots.txt. Everything else in this guide — the weekly re-audit, per-page results, the competitor crawler benchmark, Agent Readiness, Local GEO and Draft-the-fix — is the in-app version, which also tracks the score over time inside the client's project.

Two long-form guides go deeper if you want to hand a client the reasoning: "How to run a GEO audit: the complete 2026 checklist" and "GEO audit vs SEO audit", both on the CrunchJunkie blog.

Every free audit gets a permanent result page (crunchjunkie.io/products/geo-audit/r/<id>) with a share image, so a visitor can post or forward the score, the two standalone sub-scores and the fix list. Result pages are not indexed by search engines, they never show the visitor's email, and each one links straight to the free AI-visibility check for the same domain — the honest next step, because a GEO score finds blockers and does not predict how often engines name a brand.
