Log file analysis is the only way to see what AI crawlers really do on your site. robots.txt says what you allow, and analytics tools mostly miss bots, because bots rarely run the JavaScript those tools depend on. Your server logs record every request. This guide shows how to find the AI crawlers in them, how to tell real ones from fakes, and what to fix once you can see them.
What an access log line contains
Most servers write logs in the “combined” format, which both Apache and nginx define the same way: the client IP, the time, the request line, the status code, the bytes sent, the referrer and the user agent[1][2]. The last field is the one that names the crawler.
log_format combined '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'"$http_referer" "$http_user_agent"';The AI crawlers to look for
SEO log file analysis used to mean tracking Googlebot. Now it means three kinds of AI visitor as well, and they do different things. OpenAI documents four user agents with an IP list for each[3]. Anthropic runs ClaudeBot for training, Claude-User for fetches a user asks for, and Claude-SearchBot for search, all honouring robots.txt, and publishes one IP list for all three[4]. Perplexity runs PerplexityBot for its search index and Perplexity-User for user requests, which “generally ignores robots.txt rules”[5].
| Job | Follows robots.txt | Verify with | |
|---|---|---|---|
| GPTBot | Training | Yes | openai.com/gptbot.json |
| OAI-SearchBot | ChatGPT search | Yes | openai.com/searchbot.json |
| ChatGPT-User | Fetch for a user | "May not apply" | openai.com/chatgpt-user.json |
| ClaudeBot | Training | Yes | claude.com/crawling/bots.json |
| Claude-SearchBot | Claude search | Yes | claude.com/crawling/bots.json |
| Claude-User | Fetch for a user | Yes | claude.com/crawling/bots.json |
| PerplexityBot | Perplexity search | Yes | perplexity.com/perplexitybot.json |
| Perplexity-User | Fetch for a user | Generally ignores | perplexity.com/perplexity-user.json |
| CCBot | Open web corpus | Yes | index.commoncrawl.org/ccbot.json |
Common Crawl belongs on the list because it publishes its crawl as an open archive that anyone can download. Its crawler identifies itself as CCBot/2.0, and Common Crawl warns that it is “aware of crawlers falsely identifying themselves as CCBot”[8]. Google-Extended is missing from the table on purpose: it is a robots.txt token, not a crawler, so it never appears in a log.
Try it on your own log
Paste a few hundred lines from your access log. The analyser groups requests by crawler, shows the status codes each one received and the path it asked for most, and lists the check that verifies each IP. It runs entirely in your browser.
| Crawler | Job | Requests | 2xx / 3xx / 4xx / 5xx | Top path | robots.txt |
|---|---|---|---|---|---|
GPTBot OpenAI | Model training | 3 | 2 / 0 / 1 / 0 | /blog/old-post | Fetched |
OAI-SearchBot OpenAI | AI search index | 1 | 1 / 0 / 0 / 0 | /pricing | — |
ChatGPT-User OpenAI | Fetch for a user | 1 | 1 / 0 / 0 / 0 | /features | — |
Perplexity-User Perplexity | Fetch for a user | 1 | 1 / 0 / 0 / 0 | /pricing | — |
PerplexityBot Perplexity | AI search index | 1 | 0 / 0 / 0 / 1 | /docs | — |
CCBot Common Crawl | Model training | 1 | 1 / 0 / 0 / 0 | /blog | — |
How to verify these IPs
- GPTBot (192.0.2.10, 192.0.2.11): https://openai.com/gptbot.json
- OAI-SearchBot (198.51.100.7): https://openai.com/searchbot.json
- ChatGPT-User (198.51.100.9): https://openai.com/chatgpt-user.json
- Perplexity-User (203.0.113.5): https://www.perplexity.com/perplexity-user.json
- PerplexityBot (203.0.113.4): https://www.perplexity.com/perplexitybot.json
- CCBot (192.0.2.99): https://index.commoncrawl.org/ccbot.json, or reverse DNS to crawl.commoncrawl.org
Real or fake: verifying the IPs
Anyone can put GPTBot in a user agent. Cloudflare puts it plainly: user agent headers are “easily spoofed and are therefore insufficient for reliable identification”[9]. The fix is to check where the request came from. For the AI operators, that means matching the IP against their published lists. Google documents a DNS check: run a reverse lookup on the IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then run a forward lookup and confirm it returns the same IP[6]. Bing offers a Verify Bingbot tool and publishes its IPs as well[7].
host 66.249.66.1 # → crawl-66-249-66-1.googlebot.com host crawl-66-249-66-1.googlebot.com # → must return 66.249.66.1
How much fake traffic should you expect? There is no reliable industry figure, and the percentages that circulate rarely trace back to a method. One field test is worth knowing about, with its limits stated. Duane Forrester logged a new, unpromoted site for two weeks and checked every request against the published IP lists: 27 of 33 “AI assistant” requests and 692 of 799 “Googlebot” requests came from elsewhere, and some of the fakes went looking for files like .env.production[10]. He calls it one small site over 14 days, not a trend. The lesson is the method, not the number.
What to fix once you can see them
If you are deciding what to allow in the first place, should you block GPTBot works through the trade-offs, the AI crawler checker tells you what your robots.txt currently allows, and the robots.txt generator writes a new one. For the agents that act for users rather than crawl, see what replaced ChatGPT agent.
No access to raw logs?
On many managed hosts you cannot read the raw log. CDNs fill part of the gap: Cloudflare’s AI Crawl Control, generally available since August 2025, shows how AI crawlers use a site and lets paid customers answer them with a custom 402 Payment Required message instead of a plain block[11]. It is not the same as your own log, but it answers the first question: who is coming, and how often.
Questions people ask
What is log file analysis in SEO?
Reading your web server's access logs to see what crawlers actually requested: which pages, how often, and what status code they got. SEO log file analysis used to mean Googlebot; now it also shows AI crawlers like GPTBot, ClaudeBot and PerplexityBot, and the fetchers that visit when a person asks an assistant about your page.
How do I find AI crawlers in my server logs?
Filter the user-agent field for the tokens each operator documents: GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot for OpenAI; ClaudeBot, Claude-User and Claude-SearchBot for Anthropic; PerplexityBot and Perplexity-User for Perplexity; CCBot for Common Crawl. Then check the IPs against each operator's published list.
Can I see Google-Extended in my logs?
No. Google-Extended is a robots.txt token that controls how Google's AI uses your content. It is not a separate crawler, so requests arrive as Googlebot and there is no Google-Extended user agent to find.
How do I know a GPTBot request is really from OpenAI?
Check its IP against https://openai.com/gptbot.json. OpenAI publishes a separate list for each of its four user agents. A user agent alone proves nothing, because anyone can send it.
What if my host does not give me access logs?
Many CDNs and hosts offer their own bot reports instead. Cloudflare's AI Crawl Control, for example, shows AI crawler activity and lets you block individual bots. Otherwise, ask your host for raw access logs; most can export them.
