LLM tracking means asking AI engines your buyers’ questions on a schedule and recording whether they name you. The tools are the easy part. The hard part is choosing the questions: pick the wrong prompts and you will measure your brand’s fame instead of its chances of being recommended, or read noise as a trend. This post covers how to choose prompts, how many runs a number needs before you trust it, and which metrics to keep.
Why one check is not tracking
AI answers are generated fresh each time. SE Ranking ran 10,000 US keywords through Google AI Mode three times on the same day and found that “only 9.2% of URLs matched across three tests”[1]. Engines also disagree with each other: tracking 15,155 brand-category pairs daily through May 2026, Profound found a median gap of 8 percentage points in a brand’s visibility between Gemini, AI Overviews and AI Mode, and a gap of more than 10 points for over a third of brands[2].
| AI Mode | 15.2 |
|---|---|
| AI Overviews | 11.1 |
| Gemini | 6.6 |
Two things follow. Track each engine separately, because a good number on one says little about another. And measure rates across many answers, not single answers, because the same question asked twice can come back different.
Choosing the prompts
A good prompt set reads like the questions a buyer types before they have a shortlist. Five kinds cover most of the journey, plus a small branded group:
| Example | What it measures | |
|---|---|---|
| Best-of | What is the best CRM for small agencies? | Whether you make the shortlist at all |
| Problem | How can I track client deals? Which tools help? | Whether you are linked to the job, not just the category |
| Alternatives | What are the best alternatives to HubSpot? | Whether you are named when a rival is the starting point |
| Head to head | Acme CRM vs Pipedrive: which is better? | How you are described against one rival (branded) |
| Cost | How much does a CRM cost for small agencies? | Whether your pricing is known and fairly stated |
| About you | Is Acme CRM a good CRM? | What the engine believes about you (branded) |
Use your buyers’ words, not your own: the category as they say it, the jobs they are trying to get done, the rivals they already know. Sales calls, support tickets and the People Also Ask questions in your category are good sources. Then write them down and leave them alone. Add prompts when your market changes; do not swap them each month, or you lose the line you are trying to draw.
Build a starting set
Enter your category, audience, rivals and the jobs buyers hire a product like yours for. The builder writes a balanced starting set, shows how much of it names your brand, and estimates how precise a monthly mention rate would be.
- Best-ofWhat is the best CRM for small agencies?
- Best-ofRecommend a CRM for small agencies. What are the top options?
- Best-ofWhich CRM do experts recommend in 2026?
- Problem to solveHow can I track client deals? Which tools help?
- Problem to solveHow can I follow up with leads automatically? Which tools help?
- CostHow much does a CRM cost for small agencies?
- AlternativesWhat are the best alternatives to HubSpot?
- AlternativesWhat are the best alternatives to Pipedrive?
- Head to headAcme CRM vs HubSpot: which is better for small agencies?
- Head to headAcme CRM vs Pipedrive: which is better for small agencies?
- About your brandIs Acme CRM a good CRM for small agencies?
- About your brandWhat do people say about Acme CRM?
8 unbranded prompts × 4 runs = 32 answers a month per engine. A mention rate across them is accurate to about ±17.3 points (95%, worst case, treating runs as independent).
How many runs before you trust a number
A mention rate is a proportion, so ordinary sampling maths applies. At the worst case, a rate near 50%, the 95% margin of error across a month’s answers is roughly:
| Mention rate accurate to | |
|---|---|
| 1 answer | ±98.0 points |
| 10 answers | ±31.0 points |
| 40 answers | ±15.5 points |
| 100 answers | ±9.8 points |
| 400 answers | ±4.9 points |
So a single answer tells you almost nothing, and a change from 40% to 50% in a month is only real if it rests on a few hundred answers. For ±10 points you need about 97 answers; for ±5, about 385. Answers are not truly independent, since an engine may lean the same way on every run, so treat these as the best case. The practical rule: judge trends over several months, and move budget on rates, never on screenshots.
What to record for each answer
Citation and mention are different. An engine can recommend you without linking to you, and cite your guide without recommending your product. Citation analysis covers the second; the first is what most buyers act on.
Tracking it with CoreCited
CoreCited runs your prompt set every week across ChatGPT, Claude, Perplexity, Gemini, Google AI Overviews and AI Mode, and records mentions, citations, position and the rivals named instead. The Starter plan tracks 40 prompts for $69 a month. See AI visibility tracking for how scores are built and competitor tracking for share of voice.
Questions people ask
What is LLM tracking?
Asking AI engines such as ChatGPT, Gemini, Claude and Perplexity the questions your buyers ask, on a schedule, and recording whether your brand is named or cited, where it appears, and who appears instead. It is to AI answers what rank tracking is to search results.
How many prompts should I track?
Enough to cover the questions buyers ask at each stage, usually a few dozen to start. More prompts and more runs both narrow the margin of error: across 40 answers a mention rate is accurate to about ±15 points, across 400 answers to about ±5.
Should tracked prompts include my brand name?
Only a minority. A prompt that names you measures whether the engine knows you, not whether it would recommend you to someone who has never heard of you. Keep most prompts unbranded, the way a new buyer would ask.
Why do AI answers change every time I ask?
The engines generate each answer fresh and often search the web again. SE Ranking ran the same 10,000 queries through Google AI Mode three times in one day and found only 9.2% of cited URLs matched across all three runs. That is why a single check proves little.
What is AI prompt tracking?
Another name for the same practice: running a fixed set of prompts through AI engines regularly and recording the answers. The metrics that matter are mention rate, citation rate, position in the answer, share of voice against competitors, and how you are described.
