A stream of glowing machine shapes flows from the right into a neon scanning ring at the centre; cyan shapes pass through, magenta ones shatter against it, and the left side stays dark and empty.

AI Crawler Verification:Is That Really ChatGPT in Your Logs?

2026-10-01•root

Introduction

Open any web server log in 2026 and the AI names are everywhere: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Agent. Each request claims to come from a known AI company, and each claim is a line of text the client wrote about itself. Nothing stops a scraper from writing the same line.

That matters as soon as a site treats AI traffic differently from other traffic. If the firewall lets "ChatGPT" through, or the rate limiter is softer for it, the name becomes a key that anyone can copy. If the logs are used to decide which AI bots to block, a fake can get the real operator blocked. Either way the decision rests on a claim nobody checked.

To verify AI crawlers and agents, there are three checks that hold up, and each major operator supports a different mix of them. The last section turns them into a way to sort any log line into proven, fake or unverifiable.


Why the User-Agent Proves Nothing

The User-Agent header is text the client chooses. A genuine GPTBot request carries a string like Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot, and so does any script that copies it. The header was designed for compatibility and statistics, never for identity, and spoofed user agents are trivial to send.

The real identity of a request has to come from something the client cannot simply type: the network address it arrives from, a DNS record the operator controls, or a cryptographic signature made with a key only the operator holds. Those are the three methods below, and different operators support different ones.


Three Kinds of AI Traffic

An AI name in a log stands for one of three kinds of bot, and the kinds behave differently.

Training crawlers collect pages that may be used to train models. GPTBot and ClaudeBot are the main ones. Google and Apple do not run separate training crawlers: Google-Extended and Applebot-Extended are robots.txt tokens that tell the normal crawler whether its data may be used for training.

Search crawlers build the index behind an AI search feature: OAI-SearchBot for ChatGPT search, Claude-SearchBot, PerplexityBot. Googlebot and Bingbot belong here too, since their indexes feed AI answers in Google and Copilot. The common question OAI-SearchBot vs GPTBot has a short answer: blocking GPTBot keeps pages out of training, blocking OAI-SearchBot keeps them out of ChatGPT search, and the two are independent.

User-triggered fetchers and agents visit a page because a person asked for it: ChatGPT-User, Claude-User, Perplexity-User, Google-Agent, and the cloud browser behind ChatGPT's agent features. These are the ones that matter most for verification, because their robots.txt behaviour differs by operator. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, Perplexity says Perplexity-User "generally ignores" them, and Google says the same of its user-triggered fetchers. Anthropic states that its bots, Claude-User included, honour robots.txt. For the first group, robots.txt is a request at best, and the only control left is to recognise the traffic and decide what to do with it.


Method 1: Reverse DNS and Forward Confirmation

The oldest method works for classic search crawlers. Take the IP address from the log, look up its reverse DNS name, check that the name belongs to the operator's domain, then resolve that name forward and confirm it points back to the same IP. The forward step matters because anyone who controls an IP block can set its reverse DNS to googlebot.com, while only Google can make googlebot.com names resolve to that IP.

Standard DNS tools such as host, dig or nslookup do both lookups. For a genuine Googlebot address the reverse lookup returns a name like crawl-66-249-66-1.googlebot.com, and the forward lookup of that name returns the same address.

Google documents three valid domains for its crawlers: googlebot.com, google.com and googleusercontent.com. Bingbot resolves to names under search.msn.com, and Applebot to names under applebot.apple.com. A request that says Googlebot but fails this check is a fake Googlebot, and there is no softer interpretation. The reverse DNS lookup tool runs the first half of this check in a browser.

The method has a gap that matters for AI traffic. In a check on 1 October 2026, addresses from OpenAI's and Anthropic's published lists had no reverse DNS name at all, including ChatGPT-User and ClaudeBot addresses seen making real requests, and neither company documents reverse DNS as a way to verify its bots. Reverse DNS cannot verify GPTBot or ClaudeBot. For them, the published IP lists are the network-level check.


Method 2: Published IP Ranges

Most AI operators publish the address ranges their bots use, as JSON files in a common format: a creationTime and a list of ipv4Prefix and ipv6Prefix entries. A request is consistent with its claimed name if its IP falls inside one of the operator's prefixes for that bot.

OperatorBotsPublished IP list
OpenAIGPTBot, OAI-SearchBot, ChatGPT-Useropenai.com/gptbot.json, /searchbot.json, /chatgpt-user.json
AnthropicClaudeBot, Claude-SearchBot, Claude-Userclaude.com/crawling/bots.json
PerplexityPerplexityBot, Perplexity-Userperplexity.com/perplexitybot.json, /perplexity-user.json
GoogleGooglebot, special crawlers, user-triggered fetchers, Google-Agentcommon-crawlers.json, special-crawlers.json, user-triggered-fetchers.json, user-triggered-agents.json
MicrosoftBingbotbing.com/toolbox/bingbot.json
AppleApplebotsearch.developer.apple.com/applebot.json
DuckDuckGoDuckDuckBotduckduckgo.com/duckduckbot.json

Checking an address against a list is simple: download the JSON, read the prefixes, and test whether the address falls inside any of them. Most programming languages have that test in their standard library, and many firewalls and CDNs accept such lists as IP sets, so the check can run at the edge instead of in a log script.

Three details decide whether this works in practice. The lists change: on 1 October 2026 OpenAI's GPTBot file was dated 22 September and its ChatGPT-User file 25 September, so a copy fetched once and kept for months goes stale. Fetch them on a schedule and keep the last good copy if a download fails. Match the list to the bot: an address from the GPTBot list does not make a request a genuine ChatGPT-User. And remember what a match proves. The address belongs to the operator's pool for that bot, which is strong evidence for crawlers that run on the operator's own infrastructure. It says nothing about which product or user is behind the request.


Method 3: Signed Requests With Web Bot Auth

The strongest method does not depend on addresses at all. With Web Bot Auth, the bot signs each HTTP request with a private key, using the HTTP Message Signatures standard (RFC 9421). The request carries three extra headers: Signature, Signature-Input, and Signature-Agent, which names the operator's domain. The verifier fetches the operator's public keys from a fixed address on that domain, /.well-known/http-message-signatures-directory, and checks the signature. A valid signature proves the request was made by whoever holds the operator's key, from any network.

Two large operators sign today, with different coverage. OpenAI's cloud browser, the one behind ChatGPT's agent features, signs every outbound request with Signature-Agent: "https://chatgpt.com", according to OpenAI. Google signs some requests from Google-Agent with the identity https://agent.bot.goog, and calls this experimental: not every request is signed, and Google tells site owners to keep IP and reverse DNS checks as the fallback. Classic crawlers such as GPTBot and ClaudeBot do not sign.

Cloudflare and other CDNs verify these signatures for their customers and label the traffic as signed agents or verified bots, so a site behind one of them may already have the result without running a verifier itself.

A signature has limits worth knowing before trusting it fully. Web Bot Auth is still an IETF draft, and details may change. A valid signature proves who holds the key, and it does not prove that this exact request is fresh: a captured signature can be replayed until it expires, so the verifier has to enforce short validity windows. How the signature, the key directory and the replay checks fit together is covered in detail in Web Bot Auth: How the Web Started Checking an Agent's ID, and a signed request can be tested end to end with the Agent Verifier.


The AI Bot User Agent List, With How to Verify Each

User-Agent tokenWhat it doesVerify withrobots.txt
GPTBotOpenAI training crawlerIP listHonoured
OAI-SearchBotChatGPT search indexIP listHonoured
ChatGPT-UserVisits a page when a ChatGPT user asksIP listMay not apply
ChatGPT cloud browserActs on sites for a ChatGPT userSignature (chatgpt.com)User-triggered
ClaudeBotAnthropic training crawlerIP listHonoured
Claude-SearchBotClaude search qualityIP listHonoured
Claude-UserVisits a page when a Claude user asksIP listHonoured
PerplexityBotPerplexity search index, not trainingIP listHonoured
Perplexity-UserVisits a page when a Perplexity user asksIP listGenerally ignored
GooglebotGoogle Search, which also feeds AI OverviewsReverse DNS or IP listHonoured
Google-AgentAgents on Google infrastructure acting for a userIP list, some requests signed (agent.bot.goog)User-triggered
BingbotBing index, which also feeds CopilotReverse DNS or IP listHonoured
ApplebotSiri and Spotlight searchReverse DNS or IP listHonoured

"Honoured" and "may not apply" are the operators' own statements. The table describes what each operator says its bot does, and the methods above are how to check that a request with that name came from that operator at all.


How to Sort a Log Line: Proven, Fake or Unverifiable

Put the three methods in order of strength and every request that claims an AI name lands in one of three groups.

  1. Signed. If the request carries a Web Bot Auth signature, verify it against the directory of the domain in Signature-Agent. A valid signature from that operator makes the request proven. An invalid one makes it suspicious, not automatically fake, since broken signatures also come from misconfigured clients.
  2. Published IP list. If the operator publishes a list for that bot, check the address. Inside the list means the request is consistent with its name. Outside means the name is false: a GPTBot request from an address outside every OpenAI list did not come from GPTBot.
  3. Reverse DNS. For Googlebot, Bingbot and Applebot, reverse plus forward DNS gives the same answer as the list and is often easier to run on old log files.
  4. Nothing to check against. Some bots publish no list and no key. A request with such a name is unverifiable, which is different from fake. Treat it like any unknown client, and keep it out of allowlists.

The line between fake and unverifiable is easy to lose. Counting every unverifiable request as a fake inflates the problem, and counting it as genuine lets scrapers through. Keeping three outcomes, and reporting them separately, is what makes the numbers usable.

Agents that run in shared cloud browsers make the middle step weaker. When an AI product drives a browser hosted by a third-party cloud provider, its requests come from that provider's addresses, shared with every other customer. An IP list cannot tell those customers apart, and a signature can, because the key belongs to the agent operator and not to the network.


When the Bot Is Real but Still Unwelcome

Verification answers who sent a request. It does not decide whether the request should be served. For training crawlers, robots.txt remains the right tool, and the major operators honour it. For user-triggered fetchers, where robots.txt may not apply, a verified identity is what makes a precise rule possible: a rate limit for one operator, a block on one path, or a different response for agents, without touching the human visitors around them. A rule that matches only the User-Agent string also matches every scraper that copies it.


FAQ

Q: Is GPTBot safe to allow, or should it be blocked?

A: GPTBot collects pages for training OpenAI's models and honours robots.txt, so the choice is about whether the content should end up in training data, not about security. Blocking GPTBot does not remove pages from ChatGPT search, which relies on OAI-SearchBot.

Q: Does ChatGPT-User respect robots.txt?

A: Not reliably. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, because a person asked for the page. Its addresses are published, so the traffic can still be recognised and limited by IP.

Q: What is ClaudeBot, and how is it different from Claude-User?

A: ClaudeBot is Anthropic's crawler that collects web content for model training. Claude-User fetches a page when someone asks Claude about it, and Claude-SearchBot works for Claude's search. Anthropic says all three honour robots.txt and publishes one IP list for them.

Q: How can a fake Googlebot be spotted?

A: The reverse DNS name of a real Googlebot address ends in googlebot.com, google.com or googleusercontent.com, and a forward lookup of that name returns the same address. Google's published crawler ranges give the same answer. A request that claims Googlebot and fails both checks is not Googlebot.

Q: Can a scraper fake a Web Bot Auth signature?

A: Not without the operator's private key. It can copy a signature it has seen and replay it until it expires, which is why a verifier also checks the signature's time window and whether the same signature arrives twice.

Q: Why do some real AI bots fail every check?

A: Some operators publish neither IP ranges nor keys, and agents running in shared cloud browsers use addresses that belong to the cloud provider. Requests like these are unverifiable rather than fake, and the accurate result is to label them that way.


Sources


System Alert

Want to see how a server reads a signed agent request?

The Agent Verifier checks a Web Bot Auth signature, the key directory behind it and the network it arrived from, and names the exact fault when something does not match.