Somebody asks this in every marketing community, every week: how do I find out whether ChatGPT and Perplexity actually cite my website? The short answer is that you can check manually in five minutes, and the check will mislead you unless you know what you're looking at. Here's the manual method, the two traps in it, and what the question looks like when it's answered with measurement instead of a screenshot.
By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.
What citation tracking looks like when it's measured rather than spot-checked: every domain the engines cited across a prompt set, with the answers behind each count one click away.
The manual check (free, five minutes)
Ask the engines what your buyers ask. Not "tell me about {your brand}" — that's a branded prompt, and you'll pass it trivially. Ask the unbranded questions that precede a purchase in your category: best {category} for {audience}, {competitor} alternatives, is {problem} worth paying for. Then, in each answer, open the sources panel — ChatGPT lists cited links at the end of a web-grounded answer, Perplexity numbers them inline, Gemini shows them behind the answer's source chips.
Three things to record for each answer, because they are three different events: were you named in the answer text, was your domain listed as a source, and who was named or cited instead. Five prompts across two or three engines gives you a first sketch — and two systematic reasons not to trust it yet.
Trap one: the answer changes when you ask again
These models are non-deterministic and their retrieval drifts through the day. Ask the same prompt three times and you can be cited once, skipped twice — same engine, same afternoon. A screenshot of one answer is one draw from a distribution, and drawing conclusions from it is the industry's favourite mistake. We measured the effect properly in Ask Five AI Engines, Get Five Answers: the sources two identical runs cite overlap by roughly a third.
The manual check tells you the question is worth investigating. It cannot tell you a rate — for that, the same prompts have to run repeatedly, on a schedule, and the result has to carry a sample size.
Trap two: cited, retrieved and named are different wins
The second trap is vocabulary. An engine can retrieve your page (pull it in while researching), cite it (list it as a source under the answer), and name you (recommend your brand in the answer text) — and any of the three can happen without the others. Your site can be read diligently and never cited; cited for a fact while the answer recommends a competitor; or your brand named from the model's memory with your site never fetched at all. Each combination has a different fix, which is why collapsing them into one "AI visibility" number destroys the diagnosis. We took that distinction apart in Cited Isn't Named; the short version is that citations are the supply side, naming is the demand side, and you want both measured separately.
What measured looks like on a real account
Here is the same question answered with instrumentation, on our own agency's account over 90 days. The engines returned sources in a subset of runs; within the recent runs that did, our domain was cited in 90.8% ± 1.8% (n = 660) — the citation rate. In absolute terms: 996 citations across 2,112 runs, and 1,027 appearances of our pages inside answer sources across the full n = 2,848. Alongside it sits the naming side: visible in 35.6% ± 2.9% of runs. Cited far more often than named — which on this account is exactly the diagnostic worth acting on, and invisible to any single-number tool.
Measurement also buys you source health, a thing no manual check can maintain: of the 150 distinct pages the engines cited for our prompts, zero were dead links and zero had been retracted as of the last check. When an engine cites a page that later 404s, that citation is an inheritance you want to catch — someone else's dead resource is your easiest replacement opportunity.
Every scanned answer keeps its receipts: who was named, which URLs were cited. Rates are pooled from thousands of these — never from one screenshot.
The gap view: who the engines cite instead of you
Once citations are counted per domain, the most useful sort is the uncomfortable one: domains the engines cite for your prompts more than they cite you. On our account the top of that table is a competitor cited 838 times against our 301 over the window — a +537 gap on one domain — followed by two directories and an agency marketplace. That list is your outreach and content plan in ranked order. Directories and marketplaces you can join today; competitor gaps tell you which of their pages the engines trust, which is readable, specific intelligence (what do they state that we don't?).
This is also where checking turns into acting: each gap traces back to the prompts that produced it and the exact answers behind the counts, so a gap can become an evidence-grounded content brief rather than a hunch. The fan-out layer sharpens it further — the engines expose the literal searches they ran before citing anyone.
The gap table on a live account: the top row is a single competitor domain cited 838 times to our 301. Sorted this way, the citation report becomes a to-do list.
Check your own site now
The genuinely free version of everything above takes about two minutes: our AI visibility check runs live checks on your domain — no signup — and returns your visibility with the sample size and margin of error attached, who the engines recommend in your category instead, and which sources they cited. It is a snapshot and it says so; that honesty is the point.
If the snapshot bothers you, the full instrument is the same idea run on a schedule: your real buyer prompts, every engine, repeated runs, citations separated from mentions, gaps ranked, and every number carrying its n. That's the product — but start with the free check; for a surprising number of sites, the first result settles the priority argument on its own.
Frequently asked questions
Ask a question that triggers web search (most commercial questions do), then look at the end of the answer: ChatGPT lists the pages it cited as links. Perplexity numbers its citations inline; Gemini shows sources behind chips under the answer. For a real rate rather than a one-off observation, the same prompts need to run repeatedly on a schedule, because a single answer's sources change between runs.
There is no honest universal benchmark — citation rates depend entirely on your prompt set, category and engines. What you can measure is your own rate on your own buyer prompts: the share of answers-with-sources that cite your domain, with a sample size attached. On our account that is currently 90.8% ± 1.8% over 660 recent runs; treat that as an existence proof, not a target.
No. Being cited means the engine listed your page as a source; being recommended (named) means the answer told the reader to consider your brand. Either happens without the other: a page can be cited for a fact while the answer recommends a competitor, and a well-known brand can be named from the model's memory with no citation at all. Measure both separately — the gap between them is the diagnosis.
Three evidence-backed levers: make pages liftable (clear claims, sourced statistics, direct quotations — the signals shown in KDD 2024 research to raise citation likelihood by up to ~40%); be present on the third-party pages the engines already cite for your prompts, which your own citation data reveals; and make sure live-answer crawlers can actually reach your pages, since a blocked fetch silently forfeits the citation to whoever was reachable.
See your AI visibility on your own brand
Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.