---
title: "How to choose AI visibility prompts: what 2,160 AI answers showed"
url: "/blog/how-to-choose-ai-visibility-prompts"
canonical_url: "https://crunchjunkie.io/blog/how-to-choose-ai-visibility-prompts"
markdown_url: "https://crunchjunkie.io/blog/how-to-choose-ai-visibility-prompts.md"
language: "en"
last_updated: "2026-10-09"
type: "blog"
summary: "Ask an AI the same prompt twice and the answers share 55% of their brands. Reword it and they share 43%. What 2,160 answers say about choosing prompts."
site: "CrunchJunkie"
llms_txt: "https://crunchjunkie.io/llms.txt"
---

# How to choose AI visibility prompts: what 2,160 AI answers showed
Source: https://crunchjunkie.io/blog/how-to-choose-ai-visibility-prompts
Published: 2026-10-09 · By Philipp Enders · AI visibility · 22 min read

Choose the prompts you track by buyer intent, not by wording. Cover the stages where answers in your category actually name brands. Keep prompts that name a brand out of your headline number. Treat every market as its own set, and check weekly rather than daily. That is what 2,160 answers from ChatGPT, Gemini and Perplexity showed when we tested the advice most guides repeat. Those guides say to start with a few dozen prompts and to cover the funnel. They say to write each prompt several ways, add personas, mix in your brand name and reuse the set in other markets. Asking the same prompt twice changed the brands named almost as much as rewording it. Asking it a day later changed them no more than that. A worksheet and the answer-level data are free to download at the end.

## The short version
One measure runs through this post: how many brands two answers have in common. Perplexity answered the keyword prompt "sales pipeline tracking" with HubSpot, Pipedrive and Zoho CRM. A quarter of an hour later it answered HubSpot, Salesforce and Pipedrive. Between them, the two answers named four brands, and two of those were in both. So the two answers share half their brands. Identical lists share all their brands, and lists with nothing in common share none.

Asked again on the same engine, the same prompt gave answers that shared 55% of their brands on average. Two wordings of the same buyer question shared 43%, and the same prompt on two engines 39%. Most of what looks like a wording effect is already there when you simply ask again.

Waiting a day added nothing to that. One run from each day also shared 55% of their brands. None of 40 brand shares moved by more than its margin of error between the days.

Awareness prompts, about the problem a buyer has, named brands in most CRM answers. For running shoes they did so in only 55 of 216 answers, and most of those names were not shoe brands. Shortlist prompts named brands in every answer in both categories.

A prompt that names a brand gets that brand back in 83 to 100% of answers. Mixing such prompts into the set moved Brooks from second to first in the running-shoe ranking.

German answers named different brands. Across days, German and US answers shared 51% of their brands, while the two US days shared 75%. CentralStationCRM came up in more than half of the German CRM answers and in none of the 864 US ones.

How many prompts you need depends on how close the race is. Five random prompts almost always found the CRM leader. Two days of 33 prompts could not separate Asics and Brooks. For a fixed number of prompts, more intents beat more wordings.

## How we tested
We picked two categories that buyers research in very different ways. People looking for CRM software ask for tools early. People looking for running shoes often start with a sore knee.

For each category we wrote 16 buyer intents in US English, four for each of four stages. Awareness is the problem. Consideration is the shortlist. Evaluation compares named brands. For CRM the last stage is decision. For running shoes it is purchase: where to buy.

Every intent got three wordings: a plain question, a keyword string and a persona with context. "Which CRM software is free?", "free crm software" and "We're a bootstrapped startup with no budget for software. Which free CRM is good enough to start with?" are one intent in three wordings. For Germany we wrote the four consideration intents of each category in German, in the same three ways.

ChatGPT, Gemini and Perplexity answered every prompt three times on 8 October 2026. They did it all again the next day. We counted every brand an answer named, not only the ones a project tracks. Stores, apps and painkillers count too. Nine of the 32 US intents name a brand in the prompt itself, such as "hubspot vs salesforce". We report those separately.

| Design | US, per category | Germany, per category |
| --- | ---: | ---: |
| Stages | 4 | 1 (consideration) |
| Intents | 16 | 4 |
| Wordings per intent | 3 | 3 |
| Prompts | 48 | 12 |
| Answers over two days | 864 | 216 |

The study design. Every prompt ran three times on ChatGPT, Gemini and Perplexity on 8 October 2026 and again on 9 October: 1,080 answers a day, 2,160 in all.

## Asking again changes the answer almost as much as rewording it
We asked ChatGPT "What is the best CRM for a small business?" three times on 8 October. The first answer named HubSpot, Salesforce, Pipedrive, Zoho CRM and Freshsales. The second left out Salesforce. The third left out Freshsales. The keyword version, "best crm small business", also brought in monday CRM and Less Annoying CRM.

Asking again changes the list almost as much as rewording does.

Most guides tell you to track each prompt in several variations, because people phrase the same need in different ways. So we measured how much rewording changes an answer, and how much asking again does. Two runs of the same prompt on the same engine shared 55% of their brands. Two wordings of the same intent shared 43%. About four fifths of the change from rewording is already there when you just ask again.

This held on every engine and in both categories, as the table shows. The time between runs did not explain it either. Runs a few minutes apart shared about as many brands as runs 40 to 70 minutes apart.

Of the three wordings, the persona was the odd one out. The question and the keyword string were closest to each other. Personas also named fewer brands per answer.

In the marathon-shoe prompts, every answer to the plain question and to the keyword string named Nike. For the persona, a runner training for a first marathon, only 10 of 18 did. HubSpot showed the same pattern for the easiest CRM. It was in every answer to the question and the keyword string, but in 12 of 18 for a team that hated its last CRM. In effect, a persona asks a narrower question, and that question has its own answer.

So a second and third wording of one intent mostly measure the variation you get from asking again. Running each prompt more often captures that variation better. Adding intents covers more of the market.

| Answers compared | Same prompt asked again | Other wording, same intent |
| --- | ---: | ---: |
| All | 55% | 43% |
| CRM | 62% | 48% |
| Running shoes | 46% | 38% |
| ChatGPT | 64% | 52% |
| Gemini | 58% | 45% |
| Perplexity | 49% | 40% |

How many brands two answers had in common, US prompts without a brand name, both days. Each figure is an average over intents.

## Answers a day apart differed no more than answers minutes apart
On 8 October, Perplexity answered "What are the best running shoes for beginners?" with Nike, Brooks, New Balance and Altra, plus the shops Zappos and Amazon. About half an hour later, its next answer kept only Nike. The next day, its first answer was almost the same as the very first one. Only Zappos was missing.

A day between two answers changed them no more than a few minutes did.

We ran the whole study again on 9 October, about 19 hours after the first scan. If a daily check measured something real, answers from different days should differ more than answers from the same day. They did not. Two runs of one prompt shared 55% of their brands within a day, and 55% with one run from each day.

Brand shares held too. We compared 40 of them between the two days, and none moved by more than its margin of error. The largest US moves were 3.2 points, for monday CRM and Fleet Feet.

The leading brand of a prompt was no steadier. Take the brand named most often for one prompt on one engine on the first day. It still led on the second day in 124 of 185 cases. Split the same six answers into two other groups of three, mixing the days, and the leaders matched about as often. The day between the scans added almost nothing to the disagreement that asking again already produces.

So a daily AI-visibility check mostly measures how much answers vary from one run to the next. Judge a move by its margin of error, and compare weeks rather than days. We measured [how much movement is just noise](https://crunchjunkie.io/blog/ai-visibility-noise-vs-signal) over two windows in September. This test covers only two days in a row, so it says nothing about slow drift.

| What we compared | Result |
| --- | --- |
| Two runs of one prompt, both on 8 October | 55% of brands in common |
| Two runs of one prompt, both on 9 October | 55% of brands in common |
| One run from each day | 55% of brands in common |
| Brand shares that moved beyond their margin of error | 0 of 40 |
| Largest US move up | monday CRM, from 69 to 83 of 432 answers |
| Largest US move down | Fleet Feet, from 115 to 101 of 432 answers |
| Move a share near 50% needs before it counts as a change | About 7 points in the US, about 13 in Germany |
| Same leading brand for a prompt and engine on both days | 124 of 185 |
| Same leading brand when the six answers are split any other way | 64% |

Day to day, 8 and 9 October 2026. Brand shares cover the 40 brands named in at least 20 answers of a market and category on either day. The US had 432 answers per category a day, Germany 108.

## The engine matters more than the wording
On 8 October we gave all three engines the same prompt: "I run a 10-person marketing agency. Which CRM should we use?" ChatGPT's first answer named HubSpot, Salesforce, Pipedrive, Zoho CRM and Close. Gemini's named HubSpot, Pipedrive, monday CRM, GoHighLevel and Productive.io. Perplexity's named just HubSpot and Zoho CRM. Only HubSpot was in all three.

Which engine you ask changes the list more than how you word the question.

The same prompt on two different engines shared 39% of its brands. That is less than two wordings on one engine, at 43%. The engines also differ in how much they say. Gemini named the most brands per answer, Perplexity the fewest.

Choosing engines is part of choosing prompts. A number blended across three engines mixes three different lists. When the mix changes, the number moves. Keep the figures per engine next to the blend. We looked at the same effect from a brand's side in [why AI engines disagree about your brand](https://crunchjunkie.io/blog/why-ai-engines-disagree-about-your-brand).

| Engine | Brands per answer | Answers naming at least one brand |
| --- | ---: | ---: |
| Gemini | 5.44 | 554 of 576 |
| ChatGPT | 3.77 | 470 of 576 |
| Perplexity | 2.59 | 492 of 576 |

US answers in both categories over two days, 576 per engine. The same prompt on two engines shared 46% of its brands in CRM and 31% in running shoes.

## Awareness prompts name brands for CRM, rarely for running shoes
We asked "Why do my knees hurt after running?" 18 times. ChatGPT named no brand in any of its six answers. The few names that came up elsewhere were the painkillers Advil and Aleve, the running-store chain Fleet Feet and, once, the shoe brand On. Across all three wordings, 11 of 54 knee-pain answers named a brand.

Whether problem-stage prompts name brands depends on the category.

Most guides say to start at the top of the funnel, with the problem a buyer has before they know any brand. We wrote four such awareness intents per category. One CRM example: "How do small businesses keep track of sales leads?" The CRM answers named a brand in 175 of 216 cases. The running-shoe answers did so in 55 of 216.

When a shoe answer did name something, it was usually not a shoe brand. On the first day the most-named was Fleet Feet, ahead of Nike. Insoles, socks, a running app and painkillers made up much of the rest.

Further down the funnel the gap closed. Shortlist, decision and purchase prompts named brands in nearly every answer. The running-shoe purchase answers named shops about as often as shoe brands, though. On the first day Fleet Feet came up in 63 of 108 of them. Two more retailers, Road Runner Sports and Zappos, also came up often. If you make shoes, purchase prompts partly measure the shops that sell them.

Check this before you track many awareness prompts. If answers in your category rarely name brands, keep one or two to watch. Spend the rest where answers name brands.

![Grouped horizontal bar chart of US answers naming at least one brand, by stage, 216 answers per bar over two days. Awareness: CRM 175 of 216, running shoes 55 of 216. Consideration: 216 of 216 in both categories. Evaluation, where the prompt names a brand: CRM 216 of 216, running shoes 209 of 216. Decision or purchase: CRM 215 of 216, running shoes 214 of 216.](https://crunchjunkie.io/blog/img/how-to-choose-ai-visibility-prompts-stages-en.webp)

| Stage | Category | Named a brand | Brands per answer |
| --- | --- | ---: | ---: |
| Awareness (the problem) | CRM | 175 of 216 (81.0%) | 3.33 |
| Awareness (the problem) | Running shoes | 55 of 216 (25.5%) | 0.98 |
| Consideration (the shortlist) | CRM | 216 of 216 (100%) | 4.34 |
| Consideration (the shortlist) | Running shoes | 216 of 216 (100%) | 5.00 |
| Evaluation (names a brand) | CRM | 216 of 216 (100%) | 3.72 |
| Evaluation (names a brand) | Running shoes | 209 of 216 (96.8%) | 3.55 |
| Decision | CRM | 215 of 216 (99.5%) | 4.59 |
| Purchase | Running shoes | 214 of 216 (99.1%) | 5.96 |

Answers naming at least one brand, by stage, and the average number of brands per answer. US, ChatGPT, Gemini and Perplexity, three runs per prompt on each of two days: 216 answers per row. Evaluation prompts name a brand themselves.

## A prompt that names a brand measures that brand
We asked "Should a small business choose HubSpot or Salesforce?" 18 times. Every answer named both.

A brand in the prompt comes back in the answer, so the prompt mostly measures whether the engine repeats the question.

Comparisons and alternatives are popular in prompt sets: "hubspot vs salesforce", "salesforce alternatives", "Is the Nike Pegasus a good running shoe?". For most of these prompts, the named brand came back in all 54 answers. The lowest was On. Asked about the downsides of On running shoes, the answers named On in 45 of 54.

Mixed into the set, such prompts move the ranking. On the prompts without a brand name, Brooks was the second most-named running-shoe brand. Add the prompts that name Brooks and others, and Brooks came first. Asics fell from first to third, though nothing changed in how the engines treat Asics.

In CRM the order held, because HubSpot led by so much. Salesforce still gained 4.9 points from the prompts that name it.

Branded prompts are still worth tracking. They show what an engine says when someone asks about you by name. Tag them and report them apart from your visibility number.

| Brand | No brand in prompt | Rank | All prompts | Rank |
| --- | ---: | ---: | ---: | ---: |
| Asics | 49.3% (293) | 1 | 45.8% (396) | 3 |
| Brooks | 45.8% (272) | 2 | 55.7% (481) | 1 |
| Hoka | 43.4% (258) | 3 | 45.9% (397) | 2 |
| Nike | 36.9% (219) | 5 | 33.8% (292) | 5 |
| On | 5.7% (34) | 14 | 9.4% (81) | 13 |
| HubSpot | 90.4% (586) | 1 | 90.2% (779) | 1 |
| Pipedrive | 64.8% (420) | 3 | 62.2% (537) | 3 |
| Salesforce | 50.0% (324) | 4 | 54.9% (474) | 4 |

Share of US answers naming each brand, on the prompts without a brand name and on all prompts, with its rank among all brands named in the category. Running-shoe brands first (594 and 864 answers), then CRM (648 and 864). Both days.

## German and US answers name different brands
We asked for the best CRM for a small business in both markets, in English for the US and in German for Germany. The US answers mostly named HubSpot, Zoho CRM, Pipedrive and Salesforce. The German answers named HubSpot, Pipedrive and Zoho CRM too, and most of them also named CentralStationCRM. No US answer in the whole study named it.

Each market gets its own answers, with its own local brands.

For Germany we asked the same four consideration intents per category, written in German. To compare like with like, we only compared scans from different days. We set the US on one day against the US on the other day, and against Germany on the other day. Groups of three answers shared 75% of their brands between the two US days. Between a US day and a German day they shared 51%.

Some brands came up in one market only. CentralStationCRM, a German CRM, was named in 123 of 216 German CRM answers. Kalenji, a Decathlon brand, appeared only in German shoe answers. In the other direction, Less Annoying CRM and Fleet Feet appeared only in US answers.

The big names shifted as well. Microsoft Dynamics 365 and Adidas came up far more often in Germany, New Balance far less. HubSpot hardly moved.

A competitor list built for the US would have missed CentralStationCRM, the brand named in more than half of the German CRM answers.

Why compare across days? On the first day, the German re-runs were minutes apart and the US ones about 40 minutes. A same-day comparison would have mixed the two. Across days, every comparison is between two scans about 19 hours apart. Germany stands as far from the US as it did on the first day.

| Category | Brand | US | Germany |
| --- | --- | ---: | ---: |
| CRM | HubSpot | 210 (97.2%) | 205 (94.9%) |
| CRM | Zoho CRM | 160 (74.1%) | 119 (55.1%) |
| CRM | Pipedrive | 138 (63.9%) | 151 (69.9%) |
| CRM | CentralStationCRM | 0 (0.0%) | 123 (56.9%) |
| CRM | Salesforce | 95 (44.0%) | 57 (26.4%) |
| CRM | Microsoft Dynamics 365 | 15 (6.9%) | 42 (19.4%) |
| CRM | Less Annoying CRM | 46 (21.3%) | 0 (0.0%) |
| Running shoes | Asics | 194 (89.8%) | 206 (95.4%) |
| Running shoes | Brooks | 163 (75.5%) | 185 (85.6%) |
| Running shoes | New Balance | 128 (59.3%) | 71 (32.9%) |
| Running shoes | Adidas | 49 (22.7%) | 94 (43.5%) |
| Running shoes | Fleet Feet | 40 (18.5%) | 0 (0.0%) |
| Running shoes | Kalenji | 0 (0.0%) | 26 (12.0%) |

Consideration answers naming each brand, same four intents in both markets, both days. 216 answers per market and category.

## How many prompts you need depends on how close the race is
In CRM, HubSpot was named in nine of every ten answers, far ahead of Zoho CRM. In running shoes, Asics, Brooks and Hoka were each named in just under half.

A clear leader shows up with a handful of prompts. A close race may not settle even with all of them.

Guides usually give a number, often a few dozen. We tested what a given number of prompts gets right. We took the US prompts without a brand name, 36 for CRM and 33 for running shoes. For each set size we picked 500 random sets of prompts. Then we checked two things. Did a set find the same leader as all prompts together? And did it rank the brands in close to the same order?

In CRM, any five prompts found HubSpot first in 493 of 500 random sets. Any ten found it every time. The full ranking needed more prompts, and more still for the small brands at the bottom of the list. The table shows how many.

In running shoes, each of the top three had a margin of error of about four points either way. Over two days Asics edges Hoka by 5.9 points, a gap just big enough to count. Asics and Brooks, 3.5 points apart, could not be separated even with all 33 prompts. Report the two as level.

The full shoe ranking was easier to get close to than the CRM one. Ten prompts were enough, whichever brands you rank.

The winner of a single prompt is unreliable too. Take the three runs of one prompt on one engine on one day. In 300 of 414 such groups, two or more brands were tied for most named. With that many ties, who wins a single prompt is mostly chance. Read prompts in groups, by stage, intent or market.

| Category | Prompts per set | What the sets got right |
| --- | ---: | --- |
| CRM | 5 | HubSpot first in 493 of 500 sets |
| CRM | 10 | HubSpot first in all 500 sets |
| CRM | 20 | Close to the full ranking of the 21 brands named in at least ten answers |
| CRM | 25 | Close to the full ranking of all 37 brands named in at least five answers |
| Running shoes | 10 | Asics first in 310 of 500 sets, and close to the full ranking with either brand list |
| Running shoes | 20 | Asics first in 422 of 500 sets |

What random sets of US prompts without a brand name got right, 500 random sets per size, both days. "Close to the full ranking" means the set put the brands in nearly the same order as all prompts together (a rank correlation of 0.8 or better, where 1 is the same order) in nine sets out of ten.

## For the same budget, more intents beat more wordings
Say you can afford nine prompts for running shoes. You can cover nine intents in one wording each, or three intents in all three wordings. We picked 500 random sets of each kind.

Nine intents nearly always ranked the brands close to the full ranking. Three intents in three wordings got the order badly wrong about one time in ten. CRM showed the same pattern with twelve prompts, as the table shows.

This fits the first finding. Extra wordings add little beyond the variation you get from asking again. Extra intents add parts of the market you were not measuring.

| Category | Prompts | How they are spread | Typical set | Nine sets in ten at least |
| --- | ---: | --- | ---: | ---: |
| CRM | 9 | 9 intents, one wording each | 0.87 | 0.78 |
| CRM | 9 | 3 intents, three wordings each | 0.80 | 0.70 |
| CRM | 12 | 12 intents, one wording each | 0.91 | 0.85 |
| CRM | 12 | 4 intents, three wordings each | 0.83 | 0.73 |
| Running shoes | 6 | 6 intents, one wording each | 0.93 | 0.79 |
| Running shoes | 6 | 2 intents, three wordings each | 0.82 | 0.52 |
| Running shoes | 9 | 9 intents, one wording each | 0.96 | 0.90 |
| Running shoes | 9 | 3 intents, three wordings each | 0.88 | 0.66 |

How closely a smaller set of US prompts without a brand name ranked the brands like the full set, for brands named in at least ten answers (1 = the same order). 500 random sets per row: the value of the typical set, and the value nine sets in ten reached or beat. Both days.

## How to build your prompt set
1. Start from intents. Write down the jobs a buyer brings to your category at each stage. That means the problem, the shortlist, the brands they compare, and where and how they buy. One line per job. For the same number of prompts, more intents ranked the brands better than more wordings.

2. Test the top of the funnel before you commit to it. Ask each awareness candidate a few times on the engines you care about. Keep the prompts if the answers name brands. If they rarely do, keep one or two to watch and spend the rest further down the funnel. Only about one running-shoe awareness answer in four named a brand.

3. Write one wording per intent, the way your buyers ask. A plain question is a good default. Add a persona only when the buyer's situation changes what a good answer is. A tight budget, flat feet or a team that hated its last CRM are examples. Count such a persona as an intent of its own. Don't spend prompts on question and keyword versions of the same need. Two runs of one wording already share only 55% of their brands.

4. Keep brand names out of the headline. Tag every prompt that names a brand, yours or a competitor's, and report it on its own. Mixed in, such prompts moved Brooks from second to first.

5. Build each market from its own buyers. Write the prompts in the market's language, from what buyers there ask. Set the market and language for each prompt, and add local competitors to your list. No US answer named CentralStationCRM, yet more than half of the German CRM answers did.

6. Ask more than once, on more than one engine. Run each prompt several times per engine. Report the share over all runs with the number of answers behind it, and keep the figures per engine next to any blend. The same prompt on two engines shared only 39% of its brands. Which numbers to report is covered in [AI visibility metrics that matter](https://crunchjunkie.io/blog/ai-visibility-metrics-that-matter).

7. Size the set by the race you are in. Start with one prompt per intent, run it, and look at the gaps between the brands you care about. A leader far ahead shows up with a handful of prompts. Brands a few points apart may not separate at all, and then the honest report calls them level. If you need the order of the small brands too, plan for more prompts.

8. Read prompts in groups. Report by stage, intent, market and engine, not by who won prompt 14. In most groups of three answers to one prompt, two or more brands were tied for the top.

9. Check weekly, and judge every change by its margin of error. Weekly is enough to see the moves that are real. Daily data mostly shows answers varying from run to run. None of 40 brand shares moved beyond its margin of error from one day to the next.

10. Write the set down and leave it alone. Put it in the worksheet below, one row per prompt, with stage, wording and intent as tags. Then keep it stable for a few weeks. Every prompt you add or drop changes what the headline number is made of.

## Downloads
Three files, free to use under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). Credit CrunchJunkie and link to this page as the source.

[The prompt worksheet](https://crunchjunkie.io/blog/data/how-to-choose-ai-visibility-prompts-worksheet.csv) is a CSV file with one row per prompt and six example rows. Its columns are prompt, tags, category, intent, market, language and active. Tags go in one cell, separated by semicolons: "stage: consideration; wording: question; intent: free crm". The prefixes are optional.

Category takes one of four values. Use discovery for an open question about the category, brand if the prompt names your brand, competitor if it names a competitor, or shopping. The intent column means search intent and can stay blank. The buyer intents of this method live in the tags.

[The 120 study prompts](https://crunchjunkie.io/blog/data/how-to-choose-ai-visibility-prompts-study-prompts.csv) are in the same format, with our tags. The US and German sets are in one file.

[The answer-level dataset](https://crunchjunkie.io/blog/data/how-to-choose-ai-visibility-prompts-2026-10.csv) has one row per answer, 2,160 rows over the two days. Each row has the prompt with its stage, wording and intent, and whether it names a brand. It also has the engine, the day, the run, the time and the brands the answer named. It does not contain the answer texts.

The worksheet uses the import format of CrunchJunkie, the tool we build and ran this study in. There it imports under Prompts → Import, and the [setup steps are in our docs](https://crunchjunkie.io/docs/ai-visibility#section-8). A spreadsheet holds it just as well.

## What this does not show
The study covers two days in a row, 8 and 9 October 2026. Each prompt ran three times a day on three engines. Two days say nothing about drift over weeks or months. Claude, Google's AI Overviews and AI Mode, Copilot and the rest were not part of this test. Every engine changes its model without notice. We chose two categories because they differ; yours may behave like either, or like neither. Germany was tested at the consideration stage only.

The brand lists come from automatic extraction. We used them as extracted and did not re-read the 2,160 answer texts for this study. Positions in the answer were recorded only for the brand each project was set up for, HubSpot or Nike. So everything here counts whether a brand was named, not where.

The prompt-count results compare smaller sets with our own full set, which is itself only one possible set of prompts. Agreement with it is an upper bound. It does not show that a smaller set is right. Two separate sets of 10 CRM prompts agreed with each other less than either agreed with the full set.

### FAQ
**How many prompts should I track for AI visibility?**
Enough to cover the buyer intents that matter, and more where the race is close. In our test, five random CRM prompts found the clear leader in 493 of 500 random sets. Two days of 33 running-shoe prompts could not separate Asics and Brooks.

**Should I track several wordings of each prompt?**
Usually not. Two runs of the same prompt shared 55% of the brands they named, and two wordings of the same intent 43%. So most of the wording effect is already there when you ask again. Spend the budget on more intents.

**Should branded prompts count toward AI visibility?**
No. A prompt that names a brand returned that brand in 83 to 100% of answers. Mixing such prompts in moved Brooks from second to first among running-shoe brands. Track them separately.

**Can I translate my prompts for another country?**
As a start, but track the market on its own. Across days, German and US answers shared 51% of their brands, against 75% between two US days. A German CRM named in 123 of 216 German answers never appeared in the 864 US ones.

**How often should I check AI visibility?**
Weekly is enough. Answers a day apart shared 55% of their brands, the same as answers minutes apart. None of 40 brand shares moved beyond its margin of error between the days.

### Methodology & data
Measured with CrunchJunkie, which we build, on 8 and 9 October 2026: the same scan twice, about 19 hours apart. Four projects: CRM software and running shoes, each in the US (English, 48 prompts) and in Germany (German, 12 prompts). US prompts: 16 intents per category, four per stage (awareness, consideration, evaluation, and decision for CRM or purchase for running shoes), each written as a plain question, a keyword string and a persona with context. German prompts: the four consideration intents of each category in the same three wordings. Nine US intents (CRM 4, running shoes 5) name a brand and are analysed separately. Engines: ChatGPT (GPT-5.6 Luna through OpenAI's API, with web search), Gemini (3.6 Flash with Google Search grounding) and Perplexity (Agent API). Every prompt ran three times on every engine on each day: 1,080 answers a day, 2,160 in all. Within a day, runs of the same US prompt were a median of 40 minutes apart on 8 October (19 to 89) and 6 minutes on 9 October (5 to 16); German runs a median of 2 minutes. Across the two days, runs were a median of 19.4 hours apart. Twelve surplus fourth runs in the German running-shoe project on 8 October were excluded.

An answer's brand set is every brand the extraction found in it, including the project's own brand (HubSpot, Nike), after merging spelling variants. The share of brands two answers have in common is the Jaccard index of their brand sets; pairs in which neither answer named a brand are left out (counting them as identical instead does not change the picture). Re-ask and rewording pairs are formed within one day and pooled over both days: per intent, re-ask similarity averages all pairs of runs of the same prompt on the same engine, and rewording similarity all pairs of different wordings of the same intent on the same engine; the figures are means over intents, with 95% bootstrap intervals over intents (2,000 draws). The four fifths are the ratio of the two dissimilarities, (1 − re-ask) / (1 − rewording) = 0.80: asking again swapped out 45% of the brand list, rewording 57%. By pair of wordings, the question and the keyword string shared 47% of their brands, the question and the persona 42%, the keyword string and the persona 41%. Re-ask pairs 6 to 12 minutes apart on 9 October shared 52% to 56% of their brands, pairs about 40 to 70 minutes apart on 8 October 54% to 56%. Brands per answer by wording: CRM 4.16 (question), 4.22 (keyword string) and 3.88 (persona); running shoes 4.20, 4.48 and 3.38. Zoho CRM was named in 176 of 216 keyword-string answers and in 125 of 216 persona answers.

Day to day: across-day pairs combine each first-day run with each second-day run of the same prompt and engine. Brand shares (brands named in at least 20 answers of a market and category on either day, 40 in all) were compared between the days with a two-sided two-proportion test at the 5% level and with 95% Wilson intervals; none differed. A paired check over prompts excluded zero for one US brand (Zappos, −2.1 points) and three German ones, where 12 prompts per category are too few for that check. The leader of each prompt and engine on the two days was compared with the nine other ways of splitting the same six answers into two groups of three; those mixed splits agreed 64% of the time. Shares are pooled counts with 95% Wilson intervals; the margins of error quoted are those intervals.

Stages: worn-out-shoe prompts named a brand in 10 of 54 answers. In the running-shoe awareness answers of the first day, Fleet Feet was named in 19 of 108 and Nike in 12. In the running-shoe purchase answers of the first day, Fleet Feet came up in 63 of 108, Road Runner Sports in 47 and Zappos in 43. Branded prompts: the brand in the prompt came back in 54 of 54 answers for most intents, in 53 of 54 for alternatives to the Asics Gel-Kayano, in 51 of 54 for "Salesforce alternatives" in its three wordings and in 45 of 54 for the downsides of On running shoes. Among the running-shoe answers, Brooks was named in 272 of 594 answers (45.8%) on the prompts without a brand name and in 481 of 864 (55.7%) on all prompts.

US against Germany compares unions of three answers (one run of each wording) and of nine answers from scans on different days. With all nine answers of an intent on one engine pooled, the two US days shared 81% of their brands and a US day and a German day 49%.

Prompt counts: 500 random subsets per size (3, 5, 10, 15, 20, 25 and 30 prompts) of the US prompts without a brand name, each prompt with its 18 answers; Spearman rank correlation of the brands' shares against the full set and against a disjoint subset of the same size, for the brands named in at least ten answers (21 in CRM, 23 in running shoes) and in at least five (37 and 28). "Close to the full ranking" means a correlation of 0.8 or better at the 10th percentile; "nine sets in ten" is the 10th percentile and "the typical set" the median. A grid of every size gives 15 and 22 prompts for CRM and 8 and 9 for running shoes. Two disjoint sets of 10 CRM prompts matched each other at a median of 0.72 on the 21 brands. In CRM, HubSpot was named in 90.4% of these answers and Zoho CRM in 69.8%. In running shoes, Asics was named in 49.3%, Brooks in 45.8% and Hoka in 43.4% of 594 answers. The comparison of the shoe leaders uses a bootstrap over prompts of the differences between shares. The tie count (300 of 414) counts groups of three answers in which two or more brands share the top count.

The prompts and the answer-level data are published under CC BY 4.0; the answer texts are not. The headline numbers were recomputed by a second, independent script.

Dataset: https://crunchjunkie.io/blog/data/how-to-choose-ai-visibility-prompts-2026-10.csv (https://creativecommons.org/licenses/by/4.0/)

### Sources
- [Wilson, E. B. (1927). Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association 22(158)](https://doi.org/10.1080/01621459.1927.10502953)
- [Jaccard, P. (1912). The Distribution of the Flora in the Alpine Zone. New Phytologist 11(2)](https://doi.org/10.1111/j.1469-8137.1912.tb05611.x)
- [Spearman, C. (1904). The Proof and Measurement of Association between Two Things. The American Journal of Psychology 15(1)](https://doi.org/10.2307/1412159)
- [Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics 7(1)](https://doi.org/10.1214/aos/1176344552)
