AI traffic analytics shows two sides of the same relationship: how often AI crawlers such as ClaudeBot and GPTBot read your site, and how many visitors AI assistants such as ChatGPT, Claude and Perplexity send back. CoreCited puts both in one view, with the ratio between them.
Crawl-to-refer figures are Cloudflare’s cross-site measurements for the first week of August 2025[1], not CoreCited data.
What is ClaudeBot?
ClaudeBot is Anthropic’s crawler for content that could contribute to training its models. It honours robots.txt and supports Crawl-delay, and blocking it signals that your future content should be excluded from training[2]. It is one of three Anthropic crawlers: Claude-SearchBot crawls to improve search results for Claude users, and Claude-User fetches a page when someone’s question needs it. Anthropic says blocking either of those may reduce your visibility in Claude[2].
ClaudeBot shows up so often in logs because training crawlers read widely: in Cloudflare’s data, ClaudeBot and GPTBot together made up nearly half of observed AI crawling in early August 2025[1]. If that load matters to you, a Crawl-delay line slows it down; if you do not want your content used for training, block it. The Claude SEO guide has a planner for all three bots.
AI crawler directory
Every AI web crawler an operator documents, with what it does and what the operator says about robots.txt. Search it, or filter by type. Training crawlers cannot send you visitors; search and user-triggered crawlers are the ones that decide whether an assistant can cite you.
| User agent | Operator | Type | What it does | robots.txt (operator's claim) |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls content that may be used to train OpenAI's foundation models. | Controlled by robots.txt. Disallowing it means content should not be used in training. Docs → |
| OAI-SearchBot | OpenAI | Search | Surfaces sites in ChatGPT search. | Controlled by robots.txt. Opted-out sites are not shown in ChatGPT search answers, though they can still appear as navigational links. Docs → |
| ChatGPT-User | OpenAI | User-triggered | Fetches pages for user actions in ChatGPT and custom GPTs. | OpenAI says robots.txt rules "may not apply" because the actions are initiated by a user. Docs → |
| OAI-AdsBot | OpenAI | Ads review | Checks the landing pages of ads submitted to ChatGPT. Not used for training. | Only visits submitted ad landing pages. Docs → |
| ClaudeBotCrawl-delay: supported | Anthropic | Training | Collects content that could contribute to training Anthropic's models. | Honours robots.txt. Blocking it excludes future content from training. Docs → |
| Claude-SearchBotCrawl-delay: supported | Anthropic | Search | Crawls to improve search results for Claude users. | Honours robots.txt. Anthropic says blocking it may reduce visibility in Claude's search results. Docs → |
| Claude-UserCrawl-delay: supported | Anthropic | User-triggered | Fetches a page when a user's question in Claude needs it. | Honours robots.txt, with no exception stated for user requests. Docs → |
| PerplexityBot | Perplexity | Search | Surfaces and links websites in Perplexity's search results. Perplexity says it is not used to train foundation models. | Controlled by robots.txt; changes can take up to 24 hours. Docs → |
| Perplexity-User | Perplexity | User-triggered | Fetches pages in response to a user's question. | Perplexity says this fetcher "generally ignores robots.txt rules". Docs → |
| GooglebotCrawl-delay: ignored | Search | Google Search crawling. Google's AI features in Search, including AI Overviews and AI Mode, are governed by Googlebot. | Google's common crawlers always obey robots.txt when crawling automatically. Docs → | |
| Google-Extended | Control token | Not a crawler: a robots.txt token controlling whether Google may use your content to train Gemini models and for grounding in Gemini apps. | Google says it does not affect inclusion in Google Search and is not a ranking signal. Docs → | |
| Google-Agent | User-triggered | Used by agents on Google infrastructure that navigate the web and act on a user's request. | Google's user-triggered fetchers generally ignore robots.txt rules. Docs → | |
| ApplebotCrawl-delay: ignored | Apple | Search and training | Crawls for Spotlight, Siri and Safari; the data may also help train Apple's foundation models. | Respects robots.txt. With no Applebot rules, it follows your Googlebot rules. Docs → |
| Applebot-Extended | Apple | Control token | Not a crawler: a robots.txt token for opting out of training Apple's models. Disallowed pages can still appear in Apple's search features. | robots.txt token only. Docs → |
| meta-externalagent | Meta | Training | For training AI models or improving products by indexing content directly. | Controllable through robots.txt. Docs → |
| meta-externalfetcher | Meta | User-triggered | Fetches links on a user's request. | Meta says it may bypass robots.txt rules. Docs → |
| AmazonbotCrawl-delay: ignored | Amazon | Search and training | Improves Amazon's products and services; may be used to train Amazon AI models. | Honours robots.txt. Docs → |
| Amzn-SearchBot | Amazon | Search | Search experiences such as Alexa. Amazon says it does not crawl for generative AI training. | Honours robots.txt; if not named, follows the rules you give other search bots. Docs → |
| DuckAssistBot | DuckDuckGo | User-triggered | Fetches pages in real time for DuckDuckGo's AI-assisted answers. Not used to train AI models. | Honours robots.txt; a Disallow takes effect after 72 hours. Docs → |
| MistralAI-User | Mistral | User-triggered | User-initiated requests from Mistral's assistant. Not automatic crawling or training. | Controlled by robots.txt. Docs → |
| MistralAI-Training | Mistral | Training | Collects data for Mistral's training datasets. | Can be disallowed in robots.txt. Docs → |
| CCBot | Common Crawl | Training | Builds the open Common Crawl corpus, which many AI companies train on. | Honours robots.txt. Common Crawl warns that fake CCBot user agents exist. Docs → |
| Bytespider | ByteDance | Training | Reported to collect data for ByteDance products and models. | No official documentation found. Its behaviour is described only by third parties. Unverified |
Training, search and user-triggered crawlers
The same company often runs several crawlers with different jobs, and the difference decides what blocking one costs you. OpenAI, for example, runs GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User for fetches a user triggers[3].
To write rules for any of these, use the robots.txt generator and tester. To check what your live file allows today, use the AI Crawler Checker. And for the decision itself, read Should you block GPTBot?
AI referral traffic, by assistant
AI bot traffic is one side. The other is visitors, sometimes called LLM traffic. CoreCited connects to your Google Analytics 4 property and classifies sessions from ChatGPT, Perplexity, Gemini, Claude and Copilot by their referring sites, such as chatgpt.com and claude.ai. For each assistant it records the first day it sent you traffic, the closest thing there is to evidence that a visibility win turned into visitors.
For scale: across 74,752 sites in Ahrefs’ March 2026 data, AI assistants sent about 3.5 million visits, most of them from ChatGPT[6]. The guide to measuring AI traffic in GA4 shows how to set up the channel yourself.
The crawl-to-refer ratio
Put the two together and you get the crawl-to-refer ratio: pages an AI platform read for every visitor it sent. CoreCited calculates it per platform from your Cloudflare crawler data and your GA4 referrals, shows it next to the published cross-site figure, and marks it as a floor, because referrals are undercounted. Try the arithmetic with your own numbers:
How it works
What is on which plan
| Capability | From |
|---|---|
| AI referral traffic from GA4 | Starter |
| Search Console clicks and impressions | Starter |
| AI crawler analytics from Cloudflare | Growth |
| Crawl-to-refer ratio | Growth |
What it cannot tell you
- Google’s AI crawling on its own. Google’s AI features are governed by Googlebot and Google-Extended has no user agent of its own, so Google’s AI crawling cannot be separated from its search crawling[4].
- Crawlers on sites not behind Cloudflare. The crawler view reads Cloudflare’s analytics, so it needs your site to be on Cloudflare.
- Every AI visit. Visits without a referrer cannot be attributed by anyone.
Questions people ask
What is ClaudeBot?
ClaudeBot is Anthropic's web crawler for collecting content that could contribute to training its Claude models. It honours robots.txt and supports Crawl-delay. Anthropic runs two other crawlers: Claude-SearchBot for search and Claude-User for fetching pages when a user asks. Blocking ClaudeBot opts you out of training; Anthropic says blocking the other two may reduce your visibility in Claude.
Why is ClaudeBot crawling my site so much?
Training crawlers read a lot of pages, and in August 2025 Cloudflare measured Anthropic's crawl-to-refer ratio at nearly 50,000 crawls per referred visit, the highest of the major AI companies. If the load is a problem, add a Crawl-delay line for ClaudeBot, which Anthropic supports, or block it in robots.txt if you do not want your content used for training.
What is an AI crawler?
A bot that fetches web pages for an AI company. Some collect training data, some build the index an AI assistant searches, and some fetch a page at the moment a user asks about it. Operators publish their user agents so site owners can allow or block each one in robots.txt.
What is AI referral traffic?
Visits that arrive from an AI assistant, such as someone clicking a link in a ChatGPT, Claude, Perplexity or Gemini answer. Analytics tools see them as referrals from sites like chatgpt.com or claude.ai. Some visits, especially from mobile apps, arrive with no referrer and look like direct traffic, so the real number is higher than what analytics shows.
Do I need Cloudflare to see AI crawler traffic?
For CoreCited's crawler analytics, yes: it reads AI crawler requests from Cloudflare's analytics API, which works on every Cloudflare plan including free. AI referral traffic comes from Google Analytics 4 instead, and needs no Cloudflare.
Why can't I see Google's AI crawler?
Because there is not a separate one. Google's AI features in Search, including AI Overviews and AI Mode, are governed by Googlebot, and Google-Extended is a robots.txt token, not a crawler with its own user agent. So Google's AI crawling cannot be separated from its search crawling in your logs.
What is a good crawl-to-refer ratio?
Lower is better: it means an AI platform sends more visitors per page it reads. For context, Cloudflare measured about 118 crawls per referral for Perplexity, 887 for OpenAI and nearly 50,000 for Anthropic in early August 2025. Your ratio depends on your content and on how much of your AI traffic analytics can actually see.
How does CoreCited track AI referrals?
It connects to your Google Analytics 4 property and classifies sessions whose source is chatgpt.com, perplexity.ai, gemini.google.com, claude.ai or copilot.microsoft.com, among related domains, by assistant. It records the first day each assistant sent you traffic, and shows sessions per assistant over time.
