CoreCited
Feature

See which AI crawlers read your site, and which AI assistants send visitors back

Track AI crawlers like ClaudeBot and GPTBot, AI referral traffic from ChatGPT, Claude and Perplexity, and the crawl-to-refer ratio between them.

Updated 9 min read

AI traffic analytics shows two sides of the same relationship: how often AI crawlers such as ClaudeBot and GPTBot read your site, and how many visitors AI assistants such as ChatGPT, Claude and Perplexity send back. CoreCited puts both in one view, with the ratio between them.

Key takeaways
1AI crawlers come in three kinds: training, search and user-triggered. Only the last two can send you visitors.
2ClaudeBot is Anthropic's training crawler. Blocking it opts you out of training, not out of Claude's answers.
3AI referrals from ChatGPT, Claude, Perplexity, Gemini and Copilot are counted from GA4, and undercounted, because some apps send no referrer.
4The crawl-to-refer ratio shows how much a platform reads for every visitor it sends; CoreCited always reports it as a floor.
~50000
Anthropic crawls per referred visit, Aug 2025
Cloudflare
887
OpenAI crawls per referred visit
Cloudflare
118
Perplexity crawls per referred visit
Cloudflare
23
AI crawlers and tokens in the directory below
Operator docs, 2026-09-26

Crawl-to-refer figures are Cloudflare’s cross-site measurements for the first week of August 2025[1], not CoreCited data.

What is ClaudeBot?

ClaudeBot is Anthropic’s crawler for content that could contribute to training its models. It honours robots.txt and supports Crawl-delay, and blocking it signals that your future content should be excluded from training[2]. It is one of three Anthropic crawlers: Claude-SearchBot crawls to improve search results for Claude users, and Claude-User fetches a page when someone’s question needs it. Anthropic says blocking either of those may reduce your visibility in Claude[2].

ClaudeBot shows up so often in logs because training crawlers read widely: in Cloudflare’s data, ClaudeBot and GPTBot together made up nearly half of observed AI crawling in early August 2025[1]. If that load matters to you, a Crawl-delay line slows it down; if you do not want your content used for training, block it. The Claude SEO guide has a planner for all three bots.

AI crawler directory

Every AI web crawler an operator documents, with what it does and what the operator says about robots.txt. Search it, or filter by type. Training crawlers cannot send you visitors; search and user-triggered crawlers are the ones that decide whether an assistant can cite you.

User agentOperatorTypeWhat it doesrobots.txt (operator's claim)
GPTBotOpenAITrainingCrawls content that may be used to train OpenAI's foundation models.Controlled by robots.txt. Disallowing it means content should not be used in training. Docs →
OAI-SearchBotOpenAISearchSurfaces sites in ChatGPT search.Controlled by robots.txt. Opted-out sites are not shown in ChatGPT search answers, though they can still appear as navigational links. Docs →
ChatGPT-UserOpenAIUser-triggeredFetches pages for user actions in ChatGPT and custom GPTs.OpenAI says robots.txt rules "may not apply" because the actions are initiated by a user. Docs →
OAI-AdsBotOpenAIAds reviewChecks the landing pages of ads submitted to ChatGPT. Not used for training.Only visits submitted ad landing pages. Docs →
ClaudeBotCrawl-delay: supportedAnthropicTrainingCollects content that could contribute to training Anthropic's models.Honours robots.txt. Blocking it excludes future content from training. Docs →
Claude-SearchBotCrawl-delay: supportedAnthropicSearchCrawls to improve search results for Claude users.Honours robots.txt. Anthropic says blocking it may reduce visibility in Claude's search results. Docs →
Claude-UserCrawl-delay: supportedAnthropicUser-triggeredFetches a page when a user's question in Claude needs it.Honours robots.txt, with no exception stated for user requests. Docs →
PerplexityBotPerplexitySearchSurfaces and links websites in Perplexity's search results. Perplexity says it is not used to train foundation models.Controlled by robots.txt; changes can take up to 24 hours. Docs →
Perplexity-UserPerplexityUser-triggeredFetches pages in response to a user's question.Perplexity says this fetcher "generally ignores robots.txt rules". Docs →
GooglebotCrawl-delay: ignoredGoogleSearchGoogle Search crawling. Google's AI features in Search, including AI Overviews and AI Mode, are governed by Googlebot.Google's common crawlers always obey robots.txt when crawling automatically. Docs →
Google-ExtendedGoogleControl tokenNot a crawler: a robots.txt token controlling whether Google may use your content to train Gemini models and for grounding in Gemini apps.Google says it does not affect inclusion in Google Search and is not a ranking signal. Docs →
Google-AgentGoogleUser-triggeredUsed by agents on Google infrastructure that navigate the web and act on a user's request.Google's user-triggered fetchers generally ignore robots.txt rules. Docs →
ApplebotCrawl-delay: ignoredAppleSearch and trainingCrawls for Spotlight, Siri and Safari; the data may also help train Apple's foundation models.Respects robots.txt. With no Applebot rules, it follows your Googlebot rules. Docs →
Applebot-ExtendedAppleControl tokenNot a crawler: a robots.txt token for opting out of training Apple's models. Disallowed pages can still appear in Apple's search features.robots.txt token only. Docs →
meta-externalagentMetaTrainingFor training AI models or improving products by indexing content directly.Controllable through robots.txt. Docs →
meta-externalfetcherMetaUser-triggeredFetches links on a user's request.Meta says it may bypass robots.txt rules. Docs →
AmazonbotCrawl-delay: ignoredAmazonSearch and trainingImproves Amazon's products and services; may be used to train Amazon AI models.Honours robots.txt. Docs →
Amzn-SearchBotAmazonSearchSearch experiences such as Alexa. Amazon says it does not crawl for generative AI training.Honours robots.txt; if not named, follows the rules you give other search bots. Docs →
DuckAssistBotDuckDuckGoUser-triggeredFetches pages in real time for DuckDuckGo's AI-assisted answers. Not used to train AI models.Honours robots.txt; a Disallow takes effect after 72 hours. Docs →
MistralAI-UserMistralUser-triggeredUser-initiated requests from Mistral's assistant. Not automatic crawling or training.Controlled by robots.txt. Docs →
MistralAI-TrainingMistralTrainingCollects data for Mistral's training datasets.Can be disallowed in robots.txt. Docs →
CCBotCommon CrawlTrainingBuilds the open Common Crawl corpus, which many AI companies train on.Honours robots.txt. Common Crawl warns that fake CCBot user agents exist. Docs →
BytespiderByteDanceTrainingReported to collect data for ByteDance products and models.No official documentation found. Its behaviour is described only by third parties. Unverified
Each robots.txt column is what the operator says in its own documentation, read on 2026-09-26; it is not an independent test.

Training, search and user-triggered crawlers

The same company often runs several crawlers with different jobs, and the difference decides what blocking one costs you. OpenAI, for example, runs GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User for fetches a user triggers[3].

Myth
Blocking GPTBot removes you from ChatGPT's answers.
What is true
GPTBot is for training. OAI-SearchBot decides whether ChatGPT search can show your site.
Myth
Every AI crawler obeys robots.txt.
What is true
OpenAI, Perplexity, Google and Meta say some user-triggered fetchers may not follow it, because a person asked for the page.
Myth
Blocking Google-Extended hides you from Google's AI answers.
What is true
Google-Extended is a token for Gemini training and grounding. Google says it does not affect inclusion in Search, where AI Overviews and AI Mode live.

To write rules for any of these, use the robots.txt generator and tester. To check what your live file allows today, use the AI Crawler Checker. And for the decision itself, read Should you block GPTBot?

AI referral traffic, by assistant

AI bot traffic is one side. The other is visitors, sometimes called LLM traffic. CoreCited connects to your Google Analytics 4 property and classifies sessions from ChatGPT, Perplexity, Gemini, Claude and Copilot by their referring sites, such as chatgpt.com and claude.ai. For each assistant it records the first day it sent you traffic, the closest thing there is to evidence that a visibility win turned into visitors.

Why the referral count is a floor
Some AI apps open links without passing a referrer, so those visits land in Direct and no tool can attribute them. And clicks from Google’s own AI Overviews and AI Mode arrive as ordinary organic search. GA4 has had its own AI Assistant channel since May 2026, which names ChatGPT, Gemini, Deepseek, Copilot and Grok and leaves Google’s AI features in Organic Search[5]. Every AI referral number, ours included, undercounts.

For scale: across 74,752 sites in Ahrefs’ March 2026 data, AI assistants sent about 3.5 million visits, most of them from ChatGPT[6]. The guide to measuring AI traffic in GA4 shows how to set up the channel yourself.

The crawl-to-refer ratio

Put the two together and you get the crawl-to-refer ratio: pages an AI platform read for every visitor it sent. CoreCited calculates it per platform from your Cloudflare crawler data and your GA4 referrals, shows it next to the published cross-site figure, and marks it as a floor, because referrals are undercounted. Try the arithmetic with your own numbers:

Crawl-to-refer calculator
How many pages an AI platform reads for every visitor it sends you. Enter your own numbers for one platform.
300:1
pages read per visitor sent
Anthropic, cross-sitenearly 50,000:1
OpenAI, cross-site887:1
Perplexity, cross-site118:1
Benchmarks: Cloudflare, AI crawler traffic by purpose and industry, first week of August 2025. Referred visits are undercounted (some apps send no referrer), so a real ratio is no worse than the one shown.

How it works

Setting up AI traffic analytics in CoreCited
Connect GA4
Sign in with Google and pick the property. CoreCited reads sessions by source; it never writes to your analytics.
Step 1: Connect GA4. Sign in with Google and pick the property. CoreCited reads sessions by source; it never writes to your analytics.
Step 2: Connect Search Console. For clicks, impressions and positions from Google Search, alongside AI referrals.
Step 3: Connect Cloudflare. Add a read-only API token and pick the zone. AI crawler requests are matched by user agent, which works on every Cloudflare plan.
Step 4: Read the Traffic view. Crawls by bot and path, referrals by assistant with first-seen dates, and the crawl-to-refer ratio per platform.

What is on which plan

CapabilityFrom
AI referral traffic from GA4Starter
Search Console clicks and impressionsStarter
AI crawler analytics from CloudflareGrowth
Crawl-to-refer ratioGrowth

What it cannot tell you

  • Google’s AI crawling on its own. Google’s AI features are governed by Googlebot and Google-Extended has no user agent of its own, so Google’s AI crawling cannot be separated from its search crawling[4].
  • Crawlers on sites not behind Cloudflare. The crawler view reads Cloudflare’s analytics, so it needs your site to be on Cloudflare.
  • Every AI visit. Visits without a referrer cannot be attributed by anyone.

Questions people ask

What is ClaudeBot?

ClaudeBot is Anthropic's web crawler for collecting content that could contribute to training its Claude models. It honours robots.txt and supports Crawl-delay. Anthropic runs two other crawlers: Claude-SearchBot for search and Claude-User for fetching pages when a user asks. Blocking ClaudeBot opts you out of training; Anthropic says blocking the other two may reduce your visibility in Claude.

Why is ClaudeBot crawling my site so much?

Training crawlers read a lot of pages, and in August 2025 Cloudflare measured Anthropic's crawl-to-refer ratio at nearly 50,000 crawls per referred visit, the highest of the major AI companies. If the load is a problem, add a Crawl-delay line for ClaudeBot, which Anthropic supports, or block it in robots.txt if you do not want your content used for training.

What is an AI crawler?

A bot that fetches web pages for an AI company. Some collect training data, some build the index an AI assistant searches, and some fetch a page at the moment a user asks about it. Operators publish their user agents so site owners can allow or block each one in robots.txt.

What is AI referral traffic?

Visits that arrive from an AI assistant, such as someone clicking a link in a ChatGPT, Claude, Perplexity or Gemini answer. Analytics tools see them as referrals from sites like chatgpt.com or claude.ai. Some visits, especially from mobile apps, arrive with no referrer and look like direct traffic, so the real number is higher than what analytics shows.

Do I need Cloudflare to see AI crawler traffic?

For CoreCited's crawler analytics, yes: it reads AI crawler requests from Cloudflare's analytics API, which works on every Cloudflare plan including free. AI referral traffic comes from Google Analytics 4 instead, and needs no Cloudflare.

Why can't I see Google's AI crawler?

Because there is not a separate one. Google's AI features in Search, including AI Overviews and AI Mode, are governed by Googlebot, and Google-Extended is a robots.txt token, not a crawler with its own user agent. So Google's AI crawling cannot be separated from its search crawling in your logs.

What is a good crawl-to-refer ratio?

Lower is better: it means an AI platform sends more visitors per page it reads. For context, Cloudflare measured about 118 crawls per referral for Perplexity, 887 for OpenAI and nearly 50,000 for Anthropic in early August 2025. Your ratio depends on your content and on how much of your AI traffic analytics can actually see.

How does CoreCited track AI referrals?

It connects to your Google Analytics 4 property and classifies sessions whose source is chatgpt.com, perplexity.ai, gemini.google.com, claude.ai or copilot.microsoft.com, among related domains, by assistant. It records the first day each assistant sent you traffic, and shows sessions per assistant over time.

Keep reading

Sources

[1]AI crawler traffic by purpose and industry — Cloudflare, 28 August 2025
[3]Overview of OpenAI crawlers — OpenAI Developers, read 26 September 2026
[4]Google's common crawlers — Google, updated 14 July 2026
[5][GA4] Default channel group — Google Analytics Help, read 26 September 2026
[6]AI chatbot traffic — Ahrefs, March 2026

See where you stand first

Run one real question through real engines and read the answer they give. No account, no card, and the result is yours to share.