Agency rank tracking used to mean one line per client: where their pages sit in Google for the keywords that matter. That line still matters. But many buyers now read an AI answer before they see a list of links, and the answer names a handful of brands. So each client now needs a second line: whether ChatGPT, Perplexity, Gemini, Claude, AI Overviews and AI Mode name them at all. This post covers what classic tracking still does well, what it misses, how to track the second line across ten or more clients, and how to report it without promising more than the data shows.
SISTRIX tracked 82,619 prompts over 17 weeks in six countries[1]. SE Ranking ran 10,000 keywords three times in one day[3]. The survey figure comes from 602 US marketing and PR professionals; the number of agency respondents behind the 71% is not published[4].
What classic agency rank tracking still does well
A classic rank tracker checks a fixed keyword list on a schedule and records the position of each client’s best page. For agencies it does three jobs well. It shows movement early, before traffic changes. It ties work to results: a page was rewritten, and its position moved. And it scales: one account, many clients, one keyword list each, one report per client.
Positions are also relatively stable. A page that ranks fourth this week usually ranks near fourth next week, so a single weekly reading is a fair measurement. Google’s own position figure in Search Console is defined simply: the topmost position a link to your site held in the results, averaged across queries[6]. That is a number a client can check for themselves. Nothing below replaces it.
What it misses now
The trouble starts when AI answers are forced into the same model. Two examples from the vendors’ own documentation. In Search Console, “an AI Overview occupies a single position in search results, and all links in the AI Overview are assigned that same position”[6]. In Semrush Position Tracking, when a domain is cited in an AI Overview, the tool “records this as a #1 ranking”, whatever the page’s position in the regular organic results below[5]. Neither is wrong. But both mean a “position 1” in a client report can describe very different things. Our AI Overview tracking guide compares the methods in detail.
The bigger gap is outside Google. ChatGPT, Perplexity, Gemini and Claude do not show a ranked list of pages. They write an answer and name some brands. There is no position to track, only a question: did the answer name the client, and who did it name instead? A classic tracker was never built to ask that.
Why one check misleads
AI answers are generated fresh each time, and they vary a lot. SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google’s AI 2,961 times. They concluded there is “a <1 in 100 chance” that ChatGPT or Google’s AI, asked 100 times, gives the same list of brands in any two responses[2]. SE Ranking found that only 9.2% of cited URLs matched across three runs of the same query in AI Mode on one day[3].
But the variation is not total. SISTRIX tracked 82,619 prompts weekly for 17 weeks and found that “for 86% of all prompts, there is a stable core comprising a few domains; the rest rotates at a rate of 89% per week”[1]. In AI Mode that core averaged one to five domains. Their advice for projects is the right one for agencies: base “success metrics on presence over time, not on individual results”[1].
Core, emerging, carousel or absent
The stable core is the useful idea. For each prompt the client cares about, the question is not “were we named this week?” but “are we one of the few brands the answer keeps coming back to?” CoreCited answers it by counting, for each prompt and engine, the share of weekly cycles in which the client was named:
The class is not shown until there are at least 4 weekly cycles, because below that a single lucky week moves it too much. Here is what the thresholds mean in weeks:
| Core | Emerging | Carousel | Absent | |
|---|---|---|---|---|
| 4 weeks tracked | 3 to 4 | 2 to 2 | 1 to 1 | 0 |
| 8 weeks tracked | 6 to 8 | 4 to 5 | 1 to 3 | 0 |
| 12 weeks tracked | 9 to 12 | 5 to 8 | 1 to 4 | 0 |
| 17 weeks tracked | 12 to 17 | 7 to 11 | 1 to 6 | 0 |
This is more honest than a single “AI rank” for two reasons. It reads presence over time, which is what SISTRIX recommends. And it separates two very different clients: one who is named in almost every answer, and one who appeared once in a rotating slot. Both could show “mentioned” on the day you happened to check.
Tracking 10 or more clients
Running the second line across a client list is mostly a matter of discipline. Four rules do most of the work.
One prompt set per client, written once. Use the questions a buyer asks before they know the client’s name: best-of, problems to solve, alternatives to a rival, cost. Keep brand-named prompts to a small group and report them separately, because they inflate the mention rate. Our guide to choosing prompts for LLM tracking has a builder for this. Then freeze the set. Swapping prompts each month breaks the line you are trying to draw.
Weekly, not daily. The value comes from the number of answers, not how often you look. Weekly cycles are enough to classify placement after a month and to read trends over a quarter.
Each engine separately. ChatGPT, Perplexity and AI Mode draw on different sources and churn at different rates. A good rate on one says little about another.
Enough answers per client to read a rate. This is where cost and precision meet. The table divides each multi-client plan evenly across its client slots and works out the precision of a mention rate for one engine:
| Cost per client | Prompts per client | One month, one engine | Month-on-month change needed | One quarter, one engine | |
|---|---|---|---|---|---|
| Growth, 3 clients | $59.67 a month | 40 prompts | 173 answers, ±7.5 points | ±10.5 points | 520 answers, ±4.3 points |
| Agency, 10 clients | $41.90 a month | 35 prompts | 152 answers, ±7.9 points | ±11.2 points | 455 answers, ±4.6 points |
| Scale, 25 clients | $33.56 a month | 20 prompts | 87 answers, ±10.5 points | ±14.9 points | 260 answers, ±6.1 points |
Read it like this. With 35 prompts per client, a month gives about 152 answers per engine, so a mention rate is accurate to about ±7.9 points. Comparing two months, the difference has to exceed roughly ±11.2 points before you can call it real. Over a quarter the band narrows. Answers are not truly independent, since an engine can lean the same way every week, so treat these as best cases. Clients rarely need an even split, either: a large client can carry more prompts and a small one fewer.
Reporting it without overclaiming
Clients read AI visibility through an SEO lens. In the Scribewise and Scrunch survey, 80% of agencies said clients still view it that way, and 71% said they spend more time explaining GEO than doing it[4]. A report that looks like a rank report invites the wrong questions. A few rules help:
Our SEO report template for clients shows where this section fits in a monthly report, next to Search Console and GA4.
Tools, by category
The top results for this search are lists of rank trackers. We have not tested them side by side, so there is no ranking here. What matters for an agency is which category a tool belongs to, because each answers a different question.
| Measures | Good for | Limit to know | |
|---|---|---|---|
| Classic rank tracker | Google positions for a keyword list, often with SERP features | Organic movement, many clients, cheap per keyword | No view of ChatGPT, Perplexity or Claude; AI Overview scoring varies by tool |
| Search Console | Clicks, impressions and average position from Google itself | Ground truth for organic traffic | Your own sites only; AI Overview links share one position |
| AI visibility tracker | Whether AI answers name or cite a brand, per prompt and engine | The second line, share of voice against rivals | Needs repeated runs; single checks mislead |
| Reporting dashboard | Data pulled from other tools | One client-facing view | Only as good as the sources it pulls |
When you evaluate an AI visibility tracker, ask four questions. How often does it run each prompt? Does it track each engine separately? Does it show error bars or only a single score? And can a client see a report without a login?
Tracking it with CoreCited
CoreCited tracks the second line. It runs each client’s prompts weekly across ChatGPT, Claude, Perplexity, Gemini, Google AI Overviews and AI Mode, classifies every prompt as Core, emerging, carousel or absent, and shows competitor share of voice with error bars. Clients sit in one account with a brand switcher. The Agency plan covers 10 brands, 350 prompts and 5 seats for $419 a month. Public client report links start on Growth; they are not indexed, and you can revoke them or set them to expire. White-label report branding, task assignment and the read-only API start on Agency. Search Console and GA4 AI referrals connect from Starter, so the classic line can sit next to it. See CoreCited for agencies and pricing for every limit.
Questions people ask
What is agency rank tracking?
Tracking where each client's pages rank in Google for a fixed keyword list, across many clients in one account, and reporting the changes. In 2026 most agencies also need a second line per client: whether AI answers in ChatGPT, Perplexity, Gemini, Claude, AI Overviews and AI Mode name the client for its buyers' questions.
Can a normal rank tracker track ChatGPT or AI Overviews?
Partly. Some classic trackers flag AI Overviews as a search feature, and Semrush records a citation in an AI Overview as a #1 ranking. None of that tells you whether ChatGPT or Perplexity names the client, because those answers have no positions to track. That needs repeated prompts and a mention rate.
How often should an agency check AI visibility for a client?
Weekly is enough, provided you keep the same prompts. AI answers change from run to run, so the value comes from repetition, not frequency. A placement class needs at least 4 weekly cycles before it means anything.
What is the difference between Core and Carousel?
A brand is Core for a prompt when AI answers name it in at least 70% of weekly cycles, emerging at 40% to 70%, carousel when it appears only sometimes, and absent when it never appears. SISTRIX found that for 86% of prompts a stable core of a few domains persists while the rest rotate.
How do I report AI visibility to clients without overclaiming?
Report rates across many answers, never a single screenshot. Show the margin of error next to each rate, call a change significant only when it is bigger than the noise band, and report placement per prompt (Core, emerging, carousel, absent) rather than a made-up AI ranking.
