CoreCited
Technical

Log file analysis for AI crawlers: who visits, what they get, and which are fake

Log file analysis for AI crawlers: find GPTBot, ClaudeBot and PerplexityBot in your server logs, verify their IPs, and fix what they cannot read.

AlexFounder & Developer9 min read

Log file analysis is the only way to see what AI crawlers really do on your site. robots.txt says what you allow, and analytics tools mostly miss bots, because bots rarely run the JavaScript those tools depend on. Your server logs record every request. This guide shows how to find the AI crawlers in them, how to tell real ones from fakes, and what to fix once you can see them.

Key takeaways
1Every major AI operator documents its crawler names, and most publish the IP addresses they crawl from.
2A user agent is a claim. Verify it against the operator's IP list or DNS before you believe it.
3Separate training crawlers, AI search crawlers and user-triggered fetchers: they do different jobs and follow robots.txt differently.
4Look for errors. A 404 or 503 served to a search crawler is a page that cannot be cited.

What an access log line contains

Most servers write logs in the “combined” format, which both Apache and nginx define the same way: the client IP, the time, the request line, the status code, the bytes sent, the referrer and the user agent[1][2]. The last field is the one that names the crawler.

nginx: the combined log format
log_format combined '$remote_addr - $remote_user [$time_local] '
                    '"$request" $status $body_bytes_sent '
                    '"$http_referer" "$http_user_agent"';
The common format is not enough
Apache’s “common” format leaves out the referrer and user agent. If your lines end at the byte count, you cannot tell a crawler from a person. Switch to combined before you start.

The AI crawlers to look for

SEO log file analysis used to mean tracking Googlebot. Now it means three kinds of AI visitor as well, and they do different things. OpenAI documents four user agents with an IP list for each[3]. Anthropic runs ClaudeBot for training, Claude-User for fetches a user asks for, and Claude-SearchBot for search, all honouring robots.txt, and publishes one IP list for all three[4]. Perplexity runs PerplexityBot for its search index and Perplexity-User for user requests, which “generally ignores robots.txt rules”[5].

AI crawlers, and how to verify each
From each operator's own documentation, read 26 September 2026.
JobFollows robots.txtVerify with
GPTBotTrainingYesopenai.com/gptbot.json
OAI-SearchBotChatGPT searchYesopenai.com/searchbot.json
ChatGPT-UserFetch for a user"May not apply"openai.com/chatgpt-user.json
ClaudeBotTrainingYesclaude.com/crawling/bots.json
Claude-SearchBotClaude searchYesclaude.com/crawling/bots.json
Claude-UserFetch for a userYesclaude.com/crawling/bots.json
PerplexityBotPerplexity searchYesperplexity.com/perplexitybot.json
Perplexity-UserFetch for a userGenerally ignoresperplexity.com/perplexity-user.json
CCBotOpen web corpusYesindex.commoncrawl.org/ccbot.json

Common Crawl belongs on the list because it publishes its crawl as an open archive that anyone can download. Its crawler identifies itself as CCBot/2.0, and Common Crawl warns that it is “aware of crawlers falsely identifying themselves as CCBot”[8]. Google-Extended is missing from the table on purpose: it is a robots.txt token, not a crawler, so it never appears in a log.

Try it on your own log

Paste a few hundred lines from your access log. The analyser groups requests by crawler, shows the status codes each one received and the path it asked for most, and lists the check that verifies each IP. It runs entirely in your browser.

AI crawler log analyser
Paste lines from an Apache or nginx access log in the combined format. It runs in your browser; nothing is uploaded.
9 of 9 lines read · 8 from known crawlers · 1 other
CrawlerJobRequests2xx / 3xx / 4xx / 5xxTop pathrobots.txt
GPTBot
OpenAI
Model training32 / 0 / 1 / 0/blog/old-postFetched
OAI-SearchBot
OpenAI
AI search index11 / 0 / 0 / 0/pricing—
ChatGPT-User
OpenAI
Fetch for a user11 / 0 / 0 / 0/features—
Perplexity-User
Perplexity
Fetch for a user11 / 0 / 0 / 0/pricing—
PerplexityBot
Perplexity
AI search index10 / 0 / 0 / 1/docs—
CCBot
Common Crawl
Model training11 / 0 / 0 / 0/blog—
How to verify these IPs
  • GPTBot (192.0.2.10, 192.0.2.11): https://openai.com/gptbot.json
  • OAI-SearchBot (198.51.100.7): https://openai.com/searchbot.json
  • ChatGPT-User (198.51.100.9): https://openai.com/chatgpt-user.json
  • Perplexity-User (203.0.113.5): https://www.perplexity.com/perplexity-user.json
  • PerplexityBot (203.0.113.4): https://www.perplexity.com/perplexitybot.json
  • CCBot (192.0.2.99): https://index.commoncrawl.org/ccbot.json, or reverse DNS to crawl.commoncrawl.org
Grouped by the user agent each request claims. A user agent can be faked; only the IP check proves it.

Real or fake: verifying the IPs

Anyone can put GPTBot in a user agent. Cloudflare puts it plainly: user agent headers are “easily spoofed and are therefore insufficient for reliable identification”[9]. The fix is to check where the request came from. For the AI operators, that means matching the IP against their published lists. Google documents a DNS check: run a reverse lookup on the IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then run a forward lookup and confirm it returns the same IP[6]. Bing offers a Verify Bingbot tool and publishes its IPs as well[7].

Google's reverse-then-forward DNS check
host 66.249.66.1
# → crawl-66-249-66-1.googlebot.com

host crawl-66-249-66-1.googlebot.com
# → must return 66.249.66.1

How much fake traffic should you expect? There is no reliable industry figure, and the percentages that circulate rarely trace back to a method. One field test is worth knowing about, with its limits stated. Duane Forrester logged a new, unpromoted site for two weeks and checked every request against the published IP lists: 27 of 33 “AI assistant” requests and 692 of 799 “Googlebot” requests came from elsewhere, and some of the fakes went looking for files like .env.production[10]. He calls it one small site over 14 days, not a trend. The lesson is the method, not the number.

Myth
If the log says GPTBot, OpenAI crawled my site.
What is true
It says a client claimed to be GPTBot. Check the IP against openai.com/gptbot.json.
Myth
Blocking GPTBot stops ChatGPT reading my pages.
What is true
GPTBot is the training crawler. ChatGPT search uses OAI-SearchBot, and user fetches use ChatGPT-User.
Myth
Google-Extended shows up in logs.
What is true
It is a robots.txt token only. Google's requests arrive as Googlebot.

What to fix once you can see them

From log to fix
Separate the three jobs
Training, AI search and user fetches. Blocking decisions differ for each, and so does what their absence means.
Step 1: Separate the three jobs. Training, AI search and user fetches. Blocking decisions differ for each, and so does what their absence means.
Step 2: Verify before you count. Drop requests whose IP is not on the operator's list. Decisions made on fake traffic are wrong decisions.
Step 3: Find the errors. A 404, 5xx or redirect chain served to OAI-SearchBot, Claude-SearchBot or PerplexityBot is a page that cannot be cited.
Step 4: Compare with what you want read. If search crawlers never reach your pricing or product pages, check internal links, your sitemap and robots.txt.
Step 5: Repeat monthly. Crawlers change. The same checks each month show whether a fix worked.
Monthly AI crawler log check
0/6

If you are deciding what to allow in the first place, should you block GPTBot works through the trade-offs, the AI crawler checker tells you what your robots.txt currently allows, and the robots.txt generator writes a new one. For the agents that act for users rather than crawl, see what replaced ChatGPT agent.

robots.txt is what you asked for. The log is what happened. Only one of them tells you whether AI search can read the pages you want it to cite.

No access to raw logs?

On many managed hosts you cannot read the raw log. CDNs fill part of the gap: Cloudflare’s AI Crawl Control, generally available since August 2025, shows how AI crawlers use a site and lets paid customers answer them with a custom 402 Payment Required message instead of a plain block[11]. It is not the same as your own log, but it answers the first question: who is coming, and how often.

Questions people ask

What is log file analysis in SEO?

Reading your web server's access logs to see what crawlers actually requested: which pages, how often, and what status code they got. SEO log file analysis used to mean Googlebot; now it also shows AI crawlers like GPTBot, ClaudeBot and PerplexityBot, and the fetchers that visit when a person asks an assistant about your page.

How do I find AI crawlers in my server logs?

Filter the user-agent field for the tokens each operator documents: GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot for OpenAI; ClaudeBot, Claude-User and Claude-SearchBot for Anthropic; PerplexityBot and Perplexity-User for Perplexity; CCBot for Common Crawl. Then check the IPs against each operator's published list.

Can I see Google-Extended in my logs?

No. Google-Extended is a robots.txt token that controls how Google's AI uses your content. It is not a separate crawler, so requests arrive as Googlebot and there is no Google-Extended user agent to find.

How do I know a GPTBot request is really from OpenAI?

Check its IP against https://openai.com/gptbot.json. OpenAI publishes a separate list for each of its four user agents. A user agent alone proves nothing, because anyone can send it.

What if my host does not give me access logs?

Many CDNs and hosts offer their own bot reports instead. Cloudflare's AI Crawl Control, for example, shows AI crawler activity and lets you block individual bots. Otherwise, ask your host for raw access logs; most can export them.

Keep reading

Sources

[1]Log Files — Apache HTTP Server 2.4 documentation, read 26 September 2026
[2]Module ngx_http_log_module — nginx documentation, read 26 September 2026
[3]Overview of OpenAI Crawlers — OpenAI, read 26 September 2026
[5]Perplexity Crawlers — Perplexity, read 26 September 2026
[6]Verify requests from Google crawlers and fetchers — Google, read 26 September 2026
[7]Verify Bingbot — Bing Webmaster Tools, read 26 September 2026
[8]CCBot — Common Crawl, read 26 September 2026
[10]81.8% Of My 'AI Assistant' Traffic Was Fake. The Googlebot Number Was Worse — Search Engine Journal (Duane Forrester, one site), 25 June 2026
[11]Introducing AI Crawl Control — Cloudflare, 28 August 2025

Find out where you actually stand

One real question, real AI engines, and the answer they gave — including who was named in it. No account, no card.