CoreCited
Measurement

LLM tracking: choosing the prompts, and how many answers a number needs

LLM tracking done properly: which prompts to track, why most should be unbranded, how many answers a mention rate needs, and which metrics to record.

AlexFounder & Developer8 min read

LLM tracking means asking AI engines your buyers’ questions on a schedule and recording whether they name you. The tools are the easy part. The hard part is choosing the questions: pick the wrong prompts and you will measure your brand’s fame instead of its chances of being recommended, or read noise as a trend. This post covers how to choose prompts, how many runs a number needs before you trust it, and which metrics to keep.

Key takeaways
1Track the questions a buyer asks before they know your name. Keep brand-named prompts a minority.
2Cover every stage: best-of, problems to solve, alternatives, head-to-head and cost.
3AI answers change between runs, so one check is an anecdote. A rate across many answers is a measurement.
4Keep the prompt set stable. Changing the questions every month makes months impossible to compare.

Why one check is not tracking

AI answers are generated fresh each time. SE Ranking ran 10,000 US keywords through Google AI Mode three times on the same day and found that “only 9.2% of URLs matched across three tests”[1]. Engines also disagree with each other: tracking 15,155 brand-category pairs daily through May 2026, Profound found a median gap of 8 percentage points in a brand’s visibility between Gemini, AI Overviews and AI Mode, and a gap of more than 10 points for over a third of brands[2].

Citations per answer, by Google surface
Average sources cited per run, May 2026.
Citations per answer, by Google surface
AI Mode15.2
AI Overviews11.1
Gemini6.6
Profound, vendor study, July 2026

Two things follow. Track each engine separately, because a good number on one says little about another. And measure rates across many answers, not single answers, because the same question asked twice can come back different.

Choosing the prompts

A good prompt set reads like the questions a buyer types before they have a shortlist. Five kinds cover most of the journey, plus a small branded group:

The prompt mix
What each kind of prompt tells you.
ExampleWhat it measures
Best-ofWhat is the best CRM for small agencies?Whether you make the shortlist at all
ProblemHow can I track client deals? Which tools help?Whether you are linked to the job, not just the category
AlternativesWhat are the best alternatives to HubSpot?Whether you are named when a rival is the starting point
Head to headAcme CRM vs Pipedrive: which is better?How you are described against one rival (branded)
CostHow much does a CRM cost for small agencies?Whether your pricing is known and fairly stated
About youIs Acme CRM a good CRM?What the engine believes about you (branded)

Use your buyers’ words, not your own: the category as they say it, the jobs they are trying to get done, the rivals they already know. Sales calls, support tickets and the People Also Ask questions in your category are good sources. Then write them down and leave them alone. Add prompts when your market changes; do not swap them each month, or you lose the line you are trying to draw.

The branded-prompt trap
“Is Acme CRM good?” will name Acme CRM almost every time, because the question contains it. A set heavy with prompts like that reports a high mention rate that says nothing about new buyers. Keep branded prompts in their own group and report them separately.

Build a starting set

Enter your category, audience, rivals and the jobs buyers hire a product like yours for. The builder writes a balanced starting set, shows how much of it names your brand, and estimates how precise a monthly mention rate would be.

Prompt set builder
A starting list to edit, not a finished one. Separate several rivals or jobs with commas.
12 prompts · 33% name your brand
  1. Best-ofWhat is the best CRM for small agencies?
  2. Best-ofRecommend a CRM for small agencies. What are the top options?
  3. Best-ofWhich CRM do experts recommend in 2026?
  4. Problem to solveHow can I track client deals? Which tools help?
  5. Problem to solveHow can I follow up with leads automatically? Which tools help?
  6. CostHow much does a CRM cost for small agencies?
  7. AlternativesWhat are the best alternatives to HubSpot?
  8. AlternativesWhat are the best alternatives to Pipedrive?
  9. Head to headAcme CRM vs HubSpot: which is better for small agencies?
  10. Head to headAcme CRM vs Pipedrive: which is better for small agencies?
  11. About your brandIs Acme CRM a good CRM for small agencies?
  12. About your brandWhat do people say about Acme CRM?

8 unbranded prompts × 4 runs = 32 answers a month per engine. A mention rate across them is accurate to about ±17.3 points (95%, worst case, treating runs as independent).

How many runs before you trust a number

A mention rate is a proportion, so ordinary sampling maths applies. At the worst case, a rate near 50%, the 95% margin of error across a month’s answers is roughly:

Margin of error by number of answers
95% confidence, worst case (rate near 50%), treating each answer as independent. Computed on this page.
Mention rate accurate to
1 answer±98.0 points
10 answers±31.0 points
40 answers±15.5 points
100 answers±9.8 points
400 answers±4.9 points

So a single answer tells you almost nothing, and a change from 40% to 50% in a month is only real if it rests on a few hundred answers. For ±10 points you need about 97 answers; for ±5, about 385. Answers are not truly independent, since an engine may lean the same way on every run, so treat these as the best case. The practical rule: judge trends over several months, and move budget on rates, never on screenshots.

Myth
We asked ChatGPT and it named us, so we are visible.
What is true
One answer has a margin of error close to ±100 points. Ask again and it may not.
Myth
More engines is always better.
What is true
Track the engines your buyers use. Each engine needs its own sample, so spreading thin lowers precision.
Myth
A high mention rate means AI recommends us.
What is true
Only if the prompts are unbranded. Brand-named prompts inflate it.

What to record for each answer

Definition · Mention rate
The share of answers that name your brand.
Definition · Citation rate
The share of answers that link to your site as a source.
Definition · Position
Where you appear in the answer: first recommendation, one of several, or a passing mention.
Definition · Share of voice
Your mentions as a share of all brand mentions in the answers, you and rivals together.
Definition · Description
How the answer describes you, and whether it is accurate.

Citation and mention are different. An engine can recommend you without linking to you, and cite your guide without recommending your product. Citation analysis covers the second; the first is what most buyers act on.

A prompt set is a sample of your market. If it is biased towards questions you already win, the number will flatter you every month.
Before you start tracking
0/7

Tracking it with CoreCited

CoreCited runs your prompt set every week across ChatGPT, Claude, Perplexity, Gemini, Google AI Overviews and AI Mode, and records mentions, citations, position and the rivals named instead. The Starter plan tracks 40 prompts for $69 a month. See AI visibility tracking for how scores are built and competitor tracking for share of voice.

Questions people ask

What is LLM tracking?

Asking AI engines such as ChatGPT, Gemini, Claude and Perplexity the questions your buyers ask, on a schedule, and recording whether your brand is named or cited, where it appears, and who appears instead. It is to AI answers what rank tracking is to search results.

How many prompts should I track?

Enough to cover the questions buyers ask at each stage, usually a few dozen to start. More prompts and more runs both narrow the margin of error: across 40 answers a mention rate is accurate to about ±15 points, across 400 answers to about ±5.

Should tracked prompts include my brand name?

Only a minority. A prompt that names you measures whether the engine knows you, not whether it would recommend you to someone who has never heard of you. Keep most prompts unbranded, the way a new buyer would ask.

Why do AI answers change every time I ask?

The engines generate each answer fresh and often search the web again. SE Ranking ran the same 10,000 queries through Google AI Mode three times in one day and found only 9.2% of cited URLs matched across all three runs. That is why a single check proves little.

What is AI prompt tracking?

Another name for the same practice: running a fixed set of prompts through AI engines regularly and recording the answers. The metrics that matter are mention rate, citation rate, position in the answer, share of voice against competitors, and how you are described.

Keep reading

Sources

Find out where you actually stand

One real question, real AI engines, and the answer they gave — including who was named in it. No account, no card.