Is your brand visible in AI search?Start with an Audit →
Explainers10 min

How to Identify AI Crawlers in Server Logs: GPTBot, ClaudeBot, PerplexityBot and More

Learn how to identify GPTBot, ClaudeBot, PerplexityBot and other AI crawlers in server logs using user agents, IP verification, request paths and crawler purpose.

Harmeet Singh
Marketing, Publive
How to identify AI crawlers in server logs — GPTBot, ClaudeBot, PerplexityBot and moreSERVER LOGSMatch first,verifysecond.Identifying AI crawlers in server logs.DOCUMENTED USER-AGENT TOKENSGPTBotOAI-SearchBotChatGPT-UserClaudeBotClaude-SearchBotPerplexityBot

Your server log can tell you that a machine requested a page.

It can tell you when the request arrived, which URL it requested, what response your infrastructure returned and how long the request took.

What a log entry cannot automatically tell you is:

“This was a legitimate AI crawler, and this is what the request was trying to do.”

That part requires classification.

A request containing GPTBot, ClaudeBot or PerplexityBot is a useful starting signal. But the major AI platforms now operate multiple machine visitors for different purposes, and a user-agent string alone is not proof of identity.

The practical workflow is:

identify the claimed bot → understand its purpose → verify it where possible → analyse what happened

That is how raw access logs become useful AI crawler intelligence.

Start with the right log fields

You do not need a dedicated AI analytics platform to begin identifying crawler activity.

You need request-level logs with enough information preserved.

At the web-server layer,NGINX’s access-log module can record the client address, HTTP request, response status, bytes sent, user agent and request-processing time. At the CDN layer,Fastly’s logging system can expose client IP, URL path, request headers such as User-Agent, response status, response bytes and elapsed time.

For AI crawler analysis, the fields that matter most are:

FieldWhy it matters
TimestampWhen did the request happen?
Source IPWhere did it originate, and can it be verified?
HostnameWhich site or subdomain was requested?
PathWhat content did the machine request?
HTTP statusDid the request succeed?
User agentWhich crawler does the requester claim to be?
Bytes sentHow heavy was the response?
Response timeHow quickly was it served?

If the current logging setup does not preserve user agent or source IP, fix that before building an AI crawler dashboard.

Match stable crawler tokens, not complete user-agent strings

The first pass is straightforward.

Search the user-agent field for documented bot tokens.

A useful starting set is:

GPTBot
OAI-SearchBot
ChatGPT-User
ClaudeBot
Claude-SearchBot
Claude-User
PerplexityBot
Perplexity-User
Google-Agent

Do not hard-code the complete browser-like user-agent string.

For example,OpenAI’s official crawler documentation explicitly notes that version numbers in GPTBot and OAI-SearchBot user-agent strings can change. Matching the stable token such as GPTBot or OAI-SearchBot is therefore more maintainable.

This list is not meant to be an exhaustive directory of every AI crawler.

It is a reliable starting point built around major providers that publicly document both the bot name and its purpose.

OpenAI: do not combine GPTBot and ChatGPT Search traffic

One of the easiest reporting mistakes is creating a single bucket called:

OpenAI traffic

OpenAI currently documents different agents for different jobs.

According toOpenAI’s crawler reference, GPTBot crawls content that may be used in training OpenAI’s generative models, while OAI-SearchBot supports website discovery in ChatGPT Search. ChatGPT-User is used for certain user-triggered interactions and is not an automatic web crawler.

That means these three requests have different analytical meaning:

GPTBot

OAI-SearchBot

ChatGPT-User

A spike in GPTBot traffic should not be presented internally as:

ChatGPT search interest increased.

That conclusion is not supported by the request.

Keep the three categories separate.

OpenAI also publishes IP ranges for these agents, giving infrastructure teams another verification layer beyond the self-declared user agent.

Anthropic: Claude traffic also needs classification

Anthropic follows a similar pattern.

Itsofficial crawler documentation distinguishes ClaudeBot, Claude-SearchBot and Claude-User. ClaudeBot is associated with content that may contribute to model development, Claude-SearchBot supports search, and Claude-User retrieves web content at a user’s direction.

So this:

ClaudeBot

and this:

Claude-User

should not be collapsed into one metric.

They represent different types of machine behaviour.

Anthropic also publishes source-IP information that can be used to strengthen crawler verification.

Perplexity: search crawling and user retrieval are separate

Perplexity makes the same distinction between automated search crawling and user-directed activity.

Its official crawler documentation separates PerplexityBot, used to surface websites in Perplexity’s search experience, from Perplexity-User, which can fetch content when an individual user requests it.

For server-log classification, keep:

PerplexityBot

Perplexity-User

as separate values.

The useful principle is broader than Perplexity:

a search crawler and a user-triggered fetcher may come from the same company while representing very different events.

Google shows why a “contains bot” rule will fail

A simplistic crawler filter might look like:

user_agent contains "bot"

That will miss legitimate machine traffic.

Google’s user-triggered systems include Google-Agent, which is used by agents operating on Google infrastructure after a user asks them to navigate or perform an action on the web.

So your classifier should be based on a maintained provider list, not a generic assumption that automated visitors always contain bot in the user-agent name.

This will matter increasingly as machine visitors move from crawling pages to acting on behalf of users.

Classify by purpose as well as provider

Once the stable token is identified, assign a useful purpose.

A practical reporting table might look like this:

User-agent tokenProviderPurpose
GPTBotOpenAIModel development
OAI-SearchBotOpenAISearch
ChatGPT-UserOpenAIUser-triggered
ClaudeBotAnthropicModel development
Claude-SearchBotAnthropicSearch
Claude-UserAnthropicUser-triggered
PerplexityBotPerplexitySearch
Perplexity-UserPerplexityUser-triggered
Google-AgentGoogleUser-triggered agent

This is not a new industry standard or proprietary Publive framework.

It is simply a practical normalisation of the purposes documented by the providers themselves.

Cloudflare independently uses similar purpose-based labels in itsAI bot reference, separating categories such as AI Crawler, AI Search and AI Assistant.

The point is to move from:

AI requests: 180,000

to:

What kind of machine activity is actually increasing?

Treat the user-agent string as a claim, not proof

Anyone can send an HTTP request with:

User-Agent: GPTBot

That does not make the requester OpenAI.

A user-agent string is therefore useful for candidate identification, not authentication.

Google explicitly recommends verifying crawler requests using published IP ranges or reverse-DNS checks rather than relying solely on the visible user agent. Its verification method involves confirming that the source resolves to a legitimate Google hostname and then resolving that hostname back to the original IP.

OpenAI similarly publishes IP ranges alongside its crawler documentation, while Anthropic provides source-IP information for its bots.

At the infrastructure layer,Cloudflare’s Verified Bots programme uses stronger verification methods such as published IP lists, reverse DNS and Web Bot Auth rather than accepting user-agent claims alone.

The practical rule is:

match first, verify second.

The user-agent string as a claim, and the checks that verify itTHE CLAIMUser-Agent: GPTBotAnyone can send an HTTP request withthis. That does not make the requesterOpenAI.THE VERIFICATIONpublished IP rangesreverse-DNS checksverified-bot programmesmatch first, verify second.
A user-agent string is useful for candidate identification, not authentication. The practical rule is: match first, verify second.

What should the implementation actually do?

A sensible operational sequence is:

  • Capture timestamp, source IP, user agent, URL, status, bytes and response time.
  • Match the user agent against a maintained list of known AI tokens.
  • Assign the provider and documented purpose.
  • Verify the request where provider-supported verification exists.
  • Store the request as a structured crawler event.
  • Aggregate by page class, provider, purpose and response outcome.
  • Review the provider reference periodically.

This is an implementation approach, not an externally standardised seven-step framework.

The purpose is simply to turn raw log lines into data that teams can act on.

HTTP status is more useful than crawler count alone

Suppose your logs show:

8,000 OAI-SearchBot requests

That sounds substantial.

Now suppose the outcomes are:

200: 2,500

403: 3,500

429: 1,500

5xx: 500

The meaningful story is no longer:

OpenAI requested our site 8,000 times.

It is:

Most requests were unsuccessful.

Eight thousand crawler requests split by the status each one received8,000 OAI-SEARCHBOT REQUESTS2002,5004033,5004291,5005xx5002xx: retrieval succeeded · 403: access or security policy issue429: rate limiting · 5xx: origin or delivery failureA crawler dashboard without response outcomes is incomplete.
The meaningful story is no longer that OpenAI requested the site 8,000 times. It is that most requests were unsuccessful.

That is why status should be retained in every crawler event.

Standard NGINX access logs can record $status, while Fastly’s logging variables can record the final response status alongside the requested URL and client details.

A useful interpretation is:

2xx: retrieval succeeded 3xx: inspect destination and redirect behaviour 403: access or security policy issue 429: rate limiting 5xx: application or infrastructure failure

A crawler dashboard without response outcomes is incomplete.

Response size and response time reveal a second class of problem

A 200 OK does not automatically mean the page was served efficiently.

Suppose a crawler requests a product page and receives:

2 MB of scripts, markup and interface assets

to retrieve:

30 KB of meaningful content.

That does not prove the page is unusable.

But it gives you a reason to inspect the machine-facing response.

NGINX can record bytes sent and request time through fields such as $bytes_sent and $request_time.Fastly’s logging configuration can similarly include response body size, total bytes written and elapsed request time.

These fields help move the analysis from:

Did the bot reach the page?

to:

How efficiently did we serve it?

The deeper readability question is covered separately in AI Crawlability: Why a Fast Website Can Still Be Invisible to AI.

Group requests into page classes

Individual URLs matter for debugging.

Page classes are more useful for strategy.

Group requests into categories such as:

  • product
  • pricing
  • documentation
  • comparison
  • editorial
  • support
  • PDF
  • archived content

Then compare:

Which crawlers disproportionately request product pages?

Which search bots access documentation?

Where do 403 responses cluster?

Which page class has the largest payload?

Which old PDFs continue receiving search-oriented crawler traffic?

Grouping crawler requests into page classes, and the questions that unlocksGROUP REQUESTS INTO PAGE CLASSESproductpricingdocumentationcomparisoneditorialsupportPDFarchived contentTHEN COMPARE— Which crawlers disproportionately request product pages?— Which search bots access documentation?— Where do 403 responses cluster?— Which old PDFs still receive search-oriented traffic?
This is where crawler identification becomes useful beyond the infrastructure team.

This is where crawler identification becomes useful beyond the infrastructure team.

SEO, content, product and engineering teams can start looking at the same machine activity through different lenses.

Maintain a living crawler reference

AI crawler identities are evolving quickly.

New agents appear.

Existing providers introduce new purposes.

IP ranges change.

Verification methods evolve.

OpenAI already exposes separate agents for training, search and user-triggered requests. Anthropic does the same. Google is moving into user-triggered agents. Cloudflare’s bot classifications are increasingly purpose-based.

Your crawler reference should therefore live in a maintainable configuration or data table containing:

  • provider
  • stable token
  • purpose
  • verification method
  • official source
  • last-reviewed date

Do not bury it inside application logic that gets reviewed only during the next major website release.

What about bots you do not recognise?

Do not automatically classify every unusual automated user agent as an AI crawler.

Start with provider documentation.

Then use broader bot directories for discovery.

Cloudflare’sAI bot reference includes OpenAI, Anthropic and Perplexity alongside machine visitors operated by Meta, Apple, Amazon, Microsoft, ByteDance and others.

But where the provider publishes its own documentation, that should remain the stronger source for understanding what the bot actually does.

Third-party directories are useful for discovering a crawler.

The operator is the better source for its purpose.

Server logs measure access, not AI visibility

Even perfect crawler classification has a boundary.

Server logs can tell you:

  • which machine requested the page
  • which URL it requested
  • whether the request succeeded
  • how large the response was
  • how long the response took

They cannot prove that:

  • the page appeared in an AI answer
  • it was cited
  • the brand was mentioned
  • the product was recommended
  • a human later visited the website

Those are downstream events.

Crawler logging measures machine access.

It should not be presented as an AI visibility outcome.

That distinction is the reason the previous article, AI Bot Traffic: What AI Crawlers Reveal About Future Brand Discovery, treats bot traffic as an operational signal rather than customer demand.

Where Publive AXP Edge fits

Publive AXP Edge sits around this machine-request and delivery layer.

The value of crawler identification is not simply that another dashboard can say:

GPTBot visited 10,000 times.

Once relevant machine traffic is identified and classified, teams can inspect what those requests actually receive and identify access, rendering, latency or payload problems.

TheAXP Edge AI Crawler Optimization capability addresses that next layer: improving the representation served to relevant AI crawlers when the normal website response is incomplete or unnecessarily difficult for machines to consume.

The progression is:

identify → verify → classify → inspect → remediate

The first three tell you who the machine visitor is.

The last two tell you whether the website is actually serving it well.

The goal is not to know that GPTBot visited

At the most basic level, identifying AI crawler traffic is easy.

Search the user-agent field.

The useful version goes further.

Can you say:

which AI platforms are accessing the website?

what purpose each machine visitor serves?

whether the request is authentic?

which sections it accesses?

whether those requests succeed?

which page classes consistently fail?

That is when server logs become more than infrastructure exhaust.

They become an operational view of the machine audience.

Do not stop at “GPTBot visited us.”

Get to:

“We know which AI systems are requesting which parts of our site, why they are there, whether the requests are legitimate and whether our infrastructure is successfully serving them.”

That is when AI crawler logging becomes useful.

Frequently Asked Questions

How do I identify GPTBot in server logs?

Match the stable GPTBot token in the user-agent field, then use the IP information published inOpenAI’s crawler documentation when stronger verification is needed. GPTBot should remain separate from OAI-SearchBot and ChatGPT-User because OpenAI documents different purposes for each.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot is associated with content that may be used in model development, while OAI-SearchBot supports website discovery in ChatGPT Search. OpenAI gives them separate controls and published IP ranges.

How do I identify ClaudeBot traffic?

Match ClaudeBot, Claude-SearchBot and Claude-User separately.Anthropic’s official crawler guidance documents different purposes for the three and provides source-IP information for validation.

How do I identify PerplexityBot?

Match PerplexityBot separately from Perplexity-User, then use Perplexity’s provider documentation and published network information for verification where needed.

Can I trust an AI crawler user-agent string?

Not by itself. User agents can be spoofed. Provider-published IP ranges, reverse-DNS validation and increasingly cryptographic systems such asWeb Bot Auth provide stronger identity evidence.

Which log fields matter most?

At minimum, preserve timestamp, source IP, request path, HTTP status and user agent. Response size and response time are also useful. BothNGINX andFastly expose the underlying fields needed for this analysis.

Do crawler logs tell me whether my page was cited?

No. Logs prove that a machine requested a resource. Citation, brand mention, recommendation and human referral are separate downstream outcomes that require separate measurement.

AI crawlers in server logsidentify AI crawlersGPTBot trafficClaudeBot trafficPerplexityBotAI crawler user agentsAI crawler analytics
Share LinkedIn

Keep reading

All articles →

Reading about AI visibility?

See it in action.

Experience how Publive AXP makes global brands visible and recommendable — where customers actually decide: in ChatGPT, AI Overviews, Perplexity, Gemini and more.