Docs

GEO Audit

Score how well a site is positioned to be crawled, parsed and cited by AI engines.

What the GEO Audit measures

GEO (Generative Engine Optimization) is the practice of making a website easy for AI engines — ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews — to crawl, understand and cite. The GEO Audit, found under AI Visibility → GEO Audit, fetches a client's site the same way those engines do (a plain HTTP request, no JavaScript execution — which is exactly how the major AI crawlers fetch) and grades concrete, verifiable signals into a single 0–100 readiness score. Importantly, this measures readiness — whether a site is positioned to be cited — not actual citation share. How often a brand is genuinely mentioned in AI answers is tracked separately on the Visibility tab. We keep that distinction explicit so the score is never over-claimed.

The five scored categories

The composite score weights five categories. AI crawler access (30): whether your robots.txt lets the crawlers that power live AI answers — OpenAI's OAI-SearchBot and ChatGPT-User, Anthropic's Claude-User and Claude-SearchBot, PerplexityBot, Googlebot and Google-Extended — actually reach the site. Content accessibility / SSR (30): whether the page's main content is present in the raw HTML rather than injected by JavaScript, since AI crawlers don't run JS. Structured data (20): valid schema.org JSON-LD that anchors your brand as a known entity — especially an Organization block with sameAs links, plus support for richer answer-friendly types including FAQPage, Article, Product, HowTo and QAPage. Beyond just detecting these types, the audit checks that a present type actually carries its required properties (a FAQPage needs questions with accepted answers, an Article needs headline/author/datePublished, a Product needs offers, a HowTo needs steps); a type that's present but incomplete is flagged as a partial with the exact missing properties to add, rather than a full pass. Technical SEO hygiene (15): sitemap, self-referential canonical, a single H1, a sensible title and meta description, and HTTPS. llms.txt (5): presence and basic validity of an /llms.txt file — a forward-looking convention, weighted lowest and deliberately so: Google's official generative-AI guidance states it ignores llms.txt, and no answer engine has confirmed consuming it for citations, so it's scored as cheap, machine-readable hygiene and never as a Google/search ranking or citation factor. Each check is graded pass, partial or fail with a plain-language explanation of what was found and why it matters, so the result is an action list, not just a number.

Research-validated content signals

Inside the content-quality scoring, the GEO Audit checks for three writing signals that peer-reviewed research has shown make a page more likely to be quoted by generative engines: direct quotations, cited statistics, and authoritative outbound citations. These come from the KDD'24 paper "GEO: Generative Engine Optimization" (Aggarwal et al., Princeton / IIT Delhi), which measured that adding quotations, statistics, cited sources and fluent language each lifted a page's visibility in AI answers by roughly a quarter. Each check is graded pass/partial/fail with a plain-language note, and the audit labels them as research-backed so you know the recommendation isn't a hunch. The audit also grades extractable structure — whether a page breaks its content into bulleted or numbered lists and data tables, which AI answer engines lift near-verbatim into their responses, so structured content is markedly more citable than unbroken prose. It's scored fairly (a prose-rich page is nudged toward more scannable structure, never failed for it) and is grounded in 2026 research on how content structure shapes AI citation ("Structural Feature Engineering for GEO", arXiv 2603.29979). The same research identified what doesn't work: keyword stuffing was the single worst tactic, actively reducing AI visibility. So the audit treats keyword stuffing as a negative signal — it never earns points, and a page whose keyword density is too high (roughly above 6%, with an elevated band around 4–6%) is flagged as a problem to fix, not a box to tick. Nothing in the content score ever rewards keyword density; it only ever penalises stuffing.

Live AI-crawler reachability (WAF and CDN blocks)

A robots.txt rule that says "allowed" only matters if the crawler can actually reach the page. Plenty of sites unblock AI bots in robots.txt but then sit behind a WAF or CDN — Cloudflare's "Block AI bots" toggle is the common one — that quietly turns those same crawlers away at the edge. To both you and the site owner everything looks fine; to the AI engine the site is invisible. The GEO Audit now runs a live reachability check alongside the robots.txt review: it requests the page while presenting itself as each AI crawler and reports whether the request actually got through. This is informational by default — it only raises a concern when a block is genuinely confirmed, so you won't get false alarms from ordinary rate-limiting or transient errors. When a real edge block is detected, the audit names the crawler that was turned away so you can take it straight to the client's hosting or security team. The lesson it encodes: "allowed in robots.txt" and "reachable in practice" are two different things, and AI visibility needs both. You'll see this in two concrete places. On the GEO Audit, a "Firewall & access policy" card lays it out crawler by crawler — what robots.txt says next to what the live probe found, with a plain verdict for each. "Firewall blocks" is the one to act on: robots.txt allows the crawler, the edge blocks it anyway. It covers the five crawlers a block can be pinned to (ChatGPT's OAI-SearchBot and ChatGPT-User, PerplexityBot, ClaudeBot, Googlebot) and rides along into the report, shared links and the PDF. Separately, if you've wired up crawl-log ingestion, the Crawl insights page raises a WAF-block alert — mirrored by a Signal in your inbox — the moment real AI-crawler traffic starts coming back with HTTP 403/406. So you catch an edge block two ways: from a point-in-time audit, and from live traffic in between audits. One more robots.txt signal worth watching: a "Content usage policy" card appears when your site publishes Cloudflare's Content-Signal directive (search / ai-input / ai-train = yes/no). This is a preference signal, not an enforced control — no answer engine is confirmed to honour it and Google has publicly said it ignores it, so it never affects your GEO score and doesn't, on its own, change whether you're cited. We surface it for one reason: Cloudflare auto-added defaults to millions of managed robots.txt files, so your site may be broadcasting ai-input=no — asking AI engines not to use your content as live input for answers — without anyone choosing that. If you're pursuing AI visibility, that's the opposite of what you want; the card flags it (and notes when it looks like Cloudflare's default) so you can set ai-input=yes or remove the line. It's also available as the "Content usage policy" report dimension. (ai-train=no is a legitimate training opt-out with no effect on live-answer citations — the card says so, and never flags it.)

Running an audit and tracking it over time

Open a client's AI Visibility project, go to GEO Audit, and the domain is pre-filled from the client's saved website. Click Run audit and CrunchJunkie checks robots.txt, llms.txt, the sitemap, and a representative sample of pages — the homepage plus one page per site section, preferring deeper pages, so blog posts and product pages are judged, not just the shallow top level. On-page results are aggregated honestly across the sample ("3/6 sampled pages pass"), and the Per-page results section breaks the same checks down page by page, so a mixed result is always traceable to the exact page that needs fixing. Every run is saved, so you can re-audit after making changes and watch the score move — the history strip shows each snapshot with its date. The audit also re-runs automatically — weekly by default, since on-site signals move slowly and each audit reads pages from the client's own site. You can switch a project to daily re-audits in its Settings tab, and a manual Run audit is always available regardless. Automatic runs are plain page reads with no AI checks, so they never cost anything. Because the audit reflects on-site signals only, it can't see off-site factors that also drive AI visibility (brand mentions across the web, Reddit/Wikipedia presence) or whether a model was trained on your content. Treat a high score as "this site is well-positioned to be cited," and pair it with the Visibility tab to see whether that positioning is translating into actual AI mentions.

SPA and server-rendering checks

AI crawlers fetch a page with a plain HTTP request and don't run JavaScript, so anything a site paints in the browser after load is invisible to them. The audit checks for this directly. The SSR signal check measures how much real content is in the initial HTML — the words present in the raw <body> before any JavaScript runs — and passes when there's a meaningful amount, flags a partial when there's only a little, and fails on an empty shell. The client-side-only check looks for the fingerprints of a single-page-app framework (React, Next.js, Vue, Nuxt, Angular, Svelte and others) sitting on top of a near-empty initial body, which is the classic pattern of a site that renders fine for people but returns almost nothing to a crawler. The audit also checks for a robots meta noindex: if a page carries <meta name="robots" content="noindex">, it's telling search and AI answer engines to leave it out entirely, so that's called out with the exact fix. Together these three checks catch the technical reasons a page that looks perfect in a browser can still be effectively invisible to AI.

Agent readiness (emerging)

Alongside the composite GEO score, each audit now produces a separate Agent Readiness sub-score out of 100. Where the GEO score is about being crawled and cited, agent readiness is about the newer class of AI agents — the real-time assistant fetchers and tool-callers that act live on a user's behalf. It grades three concrete, controllable signals: whether robots.txt lets the real-time agent user-agents through (ChatGPT-User, Claude-User, Perplexity-User, Gemini Deep Research and similar), whether your content is readable without JavaScript (agents fetch raw HTML and usually don't run JS), and the breadth of your structured data (how many distinct schema.org types you expose for an agent to parse). Each of these carries a "Draft the fix with Crunch" button, exactly like the main audit. It is shown as a deliberately separate sub-score, not folded into your GEO score, for two reasons: it's an emerging area, and one of its signals — an agent-interface descriptor at /.well-known/mcp.json — follows a convention (MCP) that is still being standardised. That descriptor check is informational only: if we find one we flag it as forward-looking positioning and show an "MCP endpoint detected" badge, but its absence never lowers your sub-score. Nothing here penalises a site for not having adopted a pre-standard. Keeping it separate also means your historical GEO scores stay directly comparable over time.

Entity authority

Each audit also produces a separate Entity Authority sub-score out of 100 — a read on how strong your schema.org entity graph is. Where the composite score's "Structured data" category asks "is there any schema at all", entity authority asks the sharper question answer engines actually use to decide who you are and whether to trust and cite you (E-E-A-T): a complete Organization (name, url, logo and a sameAs array), links out to authoritative external profiles (LinkedIn, Crunchbase, Wikidata, official social — three or more is a strong signal), the named people behind the brand (a founder or article authors declared as Person entities), and real author attribution on your articles. It's deterministic — it only reads the structured data you actually publish, never guessing a schema type from your page text — and honest: if the sampled pages carry no article nodes, author attribution is marked "not scored" rather than counted against you (and BlogPosting counts as an Article, so a blog is never wrongly flagged as "missing Article"). Like Agent Readiness and Local GEO it's kept out of the headline GEO score, so your history stays comparable, and each gap carries a "Draft the fix with Crunch" button. You'll also see it as an "Entity authority" report dimension that rides into shared links and the PDF.

Local GEO (for local businesses)

If a client is a local business — a shop, restaurant, clinic, tradesperson, anything with a physical location or a service area — turn on "This client is a local business" in the AI Visibility → Settings tab. The GEO Audit then produces a separate Local GEO sub-score out of 100, alongside the composite GEO and Agent Readiness scores. Like Agent Readiness it's kept out of your headline GEO score — a SaaS with no LocalBusiness schema shouldn't be marked down for it — so your history stays comparable. Local GEO grades the on-site signals that make a business surface in local-intent AI answers ("best florist near me", "plumber in Berlin"): LocalBusiness schema and how complete it is, name/address/phone shown as real text and matching the schema, a link to your Google Business Profile, dedicated location pages, your city and region named in titles and headings, review markup, an embedded map, and links to reputable directories. It's honest — if you haven't set a target city it says "not evaluated" rather than failing you, and it never rewards duplicated "city-swap" location pages. Two fields sharpen the checks. Set the client's Primary city and Region so the audit can confirm you name where you operate, and paste the Google Business Profile URL — the entity link that ties your site to your Maps listing, one of the strongest local signals. To find that URL: open Google Maps, search the business, open its listing, click Share and copy the link (it looks like https://maps.app.goo.gl/… or https://www.google.com/maps/place/…). Then run the audit to see the Local GEO card, with a "Draft the fix with Crunch" button on every issue just like the main audit.

Draft the fix with Crunch

Every issue the audit raises comes with a plain-language explanation of what was found and how to change it. For the ones where you'd rather not write the fix from scratch, each issue has a "Draft the fix with Crunch" button. Click it and Crunch drafts a concrete, copy-pasteable remediation for that specific issue — grounded in the actual finding and the audited page, not generic advice — with a Copy button to lift it straight into a ticket or hand it to the client's developer. The drafted fix is generated on demand for the issue in front of you and isn't saved to the audit; re-run it any time. Like the rest of Crunch, it needs an available AI model (Managed AI or your own key). It turns the audit from a list of problems into a list of ready-to-apply changes. One issue has its own dedicated button. When the /llms.txt check isn't passing, the llms.txt category shows Generate llms.txt with Crunch: it drafts a complete starter /llms.txt from the site's OWN structure — the real URLs in its sitemap.xml plus the homepage title and description — grouped into sections with a short description per link. Every link is a real page on the site, never an invented one; save the result at your domain root as /llms.txt after a quick review.

Share it: the free public audit

The same audit engine also powers a free, public version at crunchjunkie.io/products/geo-audit — no login, no account. Anyone can enter a domain and get the 0–100 score, the category breakdown and a prioritised fix list in seconds. It's handy for agencies: run a quick GEO check on a prospect's site before a pitch, or send the link so a client can see the problem for themselves. The public tool is rate-limited and reads only public pages and robots.txt. Everything else in this guide — the weekly re-audit, per-page results, the competitor crawler benchmark, Agent Readiness, Local GEO and Draft-the-fix — is the in-app version, which also tracks the score over time inside the client's project. Two long-form guides go deeper if you want to hand a client the reasoning: "How to run a GEO audit: the complete 2026 checklist" and "GEO audit vs SEO audit", both on the CrunchJunkie blog.