Your server log can tell you that a machine requested a page.
It can tell you when the request arrived, which URL it requested, what response your infrastructure returned and how long the request took.
What a log entry cannot automatically tell you is:
“This was a legitimate AI crawler, and this is what the request was trying to do.”
That part requires classification.
A request containing GPTBot, ClaudeBot or PerplexityBot is a useful starting signal. But the major AI platforms now operate multiple machine visitors for different purposes, and a user-agent string alone is not proof of identity.
The practical workflow is:
identify the claimed bot → understand its purpose → verify it where possible → analyse what happened
That is how raw access logs become useful AI crawler intelligence.
Start with the right log fields
You do not need a dedicated AI analytics platform to begin identifying crawler activity.
You need request-level logs with enough information preserved.
At the web-server layer,NGINX’s access-log module can record the client address, HTTP request, response status, bytes sent, user agent and request-processing time. At the CDN layer,Fastly’s logging system can expose client IP, URL path, request headers such as User-Agent, response status, response bytes and elapsed time.
For AI crawler analysis, the fields that matter most are:
| Field | Why it matters |
|---|---|
| Timestamp | When did the request happen? |
| Source IP | Where did it originate, and can it be verified? |
| Hostname | Which site or subdomain was requested? |
| Path | What content did the machine request? |
| HTTP status | Did the request succeed? |
| User agent | Which crawler does the requester claim to be? |
| Bytes sent | How heavy was the response? |
| Response time | How quickly was it served? |
If the current logging setup does not preserve user agent or source IP, fix that before building an AI crawler dashboard.
Match stable crawler tokens, not complete user-agent strings
The first pass is straightforward.
Search the user-agent field for documented bot tokens.
A useful starting set is:
GPTBot
OAI-SearchBot
ChatGPT-User
ClaudeBot
Claude-SearchBot
Claude-User
PerplexityBot
Perplexity-User
Google-AgentDo not hard-code the complete browser-like user-agent string.
For example,OpenAI’s official crawler documentation explicitly notes that version numbers in GPTBot and OAI-SearchBot user-agent strings can change. Matching the stable token such as GPTBot or OAI-SearchBot is therefore more maintainable.
This list is not meant to be an exhaustive directory of every AI crawler.
It is a reliable starting point built around major providers that publicly document both the bot name and its purpose.
OpenAI: do not combine GPTBot and ChatGPT Search traffic
One of the easiest reporting mistakes is creating a single bucket called:
OpenAI traffic
OpenAI currently documents different agents for different jobs.
According toOpenAI’s crawler reference, GPTBot crawls content that may be used in training OpenAI’s generative models, while OAI-SearchBot supports website discovery in ChatGPT Search. ChatGPT-User is used for certain user-triggered interactions and is not an automatic web crawler.
That means these three requests have different analytical meaning:
GPTBot
OAI-SearchBot
ChatGPT-User
A spike in GPTBot traffic should not be presented internally as:
ChatGPT search interest increased.
That conclusion is not supported by the request.
Keep the three categories separate.
OpenAI also publishes IP ranges for these agents, giving infrastructure teams another verification layer beyond the self-declared user agent.
Anthropic: Claude traffic also needs classification
Anthropic follows a similar pattern.
Itsofficial crawler documentation distinguishes ClaudeBot, Claude-SearchBot and Claude-User. ClaudeBot is associated with content that may contribute to model development, Claude-SearchBot supports search, and Claude-User retrieves web content at a user’s direction.
So this:
ClaudeBot
and this:
Claude-User
should not be collapsed into one metric.
They represent different types of machine behaviour.
Anthropic also publishes source-IP information that can be used to strengthen crawler verification.
Perplexity: search crawling and user retrieval are separate
Perplexity makes the same distinction between automated search crawling and user-directed activity.
Its official crawler documentation separates PerplexityBot, used to surface websites in Perplexity’s search experience, from Perplexity-User, which can fetch content when an individual user requests it.
For server-log classification, keep:
PerplexityBot
Perplexity-User
as separate values.
The useful principle is broader than Perplexity:
a search crawler and a user-triggered fetcher may come from the same company while representing very different events.
Google shows why a “contains bot” rule will fail
A simplistic crawler filter might look like:
user_agent contains "bot"
That will miss legitimate machine traffic.
Google’s user-triggered systems include Google-Agent, which is used by agents operating on Google infrastructure after a user asks them to navigate or perform an action on the web.
So your classifier should be based on a maintained provider list, not a generic assumption that automated visitors always contain bot in the user-agent name.
This will matter increasingly as machine visitors move from crawling pages to acting on behalf of users.
Classify by purpose as well as provider
Once the stable token is identified, assign a useful purpose.
A practical reporting table might look like this:
| User-agent token | Provider | Purpose |
|---|---|---|
| GPTBot | OpenAI | Model development |
| OAI-SearchBot | OpenAI | Search |
| ChatGPT-User | OpenAI | User-triggered |
| ClaudeBot | Anthropic | Model development |
| Claude-SearchBot | Anthropic | Search |
| Claude-User | Anthropic | User-triggered |
| PerplexityBot | Perplexity | Search |
| Perplexity-User | Perplexity | User-triggered |
| Google-Agent | User-triggered agent |
This is not a new industry standard or proprietary Publive framework.
It is simply a practical normalisation of the purposes documented by the providers themselves.
Cloudflare independently uses similar purpose-based labels in itsAI bot reference, separating categories such as AI Crawler, AI Search and AI Assistant.
The point is to move from:
AI requests: 180,000
to:
What kind of machine activity is actually increasing?
Treat the user-agent string as a claim, not proof
Anyone can send an HTTP request with:
User-Agent: GPTBotThat does not make the requester OpenAI.
A user-agent string is therefore useful for candidate identification, not authentication.
Google explicitly recommends verifying crawler requests using published IP ranges or reverse-DNS checks rather than relying solely on the visible user agent. Its verification method involves confirming that the source resolves to a legitimate Google hostname and then resolving that hostname back to the original IP.
OpenAI similarly publishes IP ranges alongside its crawler documentation, while Anthropic provides source-IP information for its bots.
At the infrastructure layer,Cloudflare’s Verified Bots programme uses stronger verification methods such as published IP lists, reverse DNS and Web Bot Auth rather than accepting user-agent claims alone.
The practical rule is:
match first, verify second.
What should the implementation actually do?
A sensible operational sequence is:
- Capture timestamp, source IP, user agent, URL, status, bytes and response time.
- Match the user agent against a maintained list of known AI tokens.
- Assign the provider and documented purpose.
- Verify the request where provider-supported verification exists.
- Store the request as a structured crawler event.
- Aggregate by page class, provider, purpose and response outcome.
- Review the provider reference periodically.
This is an implementation approach, not an externally standardised seven-step framework.
The purpose is simply to turn raw log lines into data that teams can act on.
HTTP status is more useful than crawler count alone
Suppose your logs show:
8,000 OAI-SearchBot requests
That sounds substantial.
Now suppose the outcomes are:
200: 2,500
403: 3,500
429: 1,500
5xx: 500
The meaningful story is no longer:
OpenAI requested our site 8,000 times.
It is:
Most requests were unsuccessful.
That is why status should be retained in every crawler event.
Standard NGINX access logs can record $status, while Fastly’s logging variables can record the final response status alongside the requested URL and client details.
A useful interpretation is:
2xx: retrieval succeeded 3xx: inspect destination and redirect behaviour 403: access or security policy issue 429: rate limiting 5xx: application or infrastructure failure
A crawler dashboard without response outcomes is incomplete.
Response size and response time reveal a second class of problem
A 200 OK does not automatically mean the page was served efficiently.
Suppose a crawler requests a product page and receives:
2 MB of scripts, markup and interface assets
to retrieve:
30 KB of meaningful content.
That does not prove the page is unusable.
But it gives you a reason to inspect the machine-facing response.
NGINX can record bytes sent and request time through fields such as $bytes_sent and $request_time.Fastly’s logging configuration can similarly include response body size, total bytes written and elapsed request time.
These fields help move the analysis from:
Did the bot reach the page?
to:
How efficiently did we serve it?
The deeper readability question is covered separately in AI Crawlability: Why a Fast Website Can Still Be Invisible to AI.
Group requests into page classes
Individual URLs matter for debugging.
Page classes are more useful for strategy.
Group requests into categories such as:
- product
- pricing
- documentation
- comparison
- editorial
- support
- archived content
Then compare:
Which crawlers disproportionately request product pages?
Which search bots access documentation?
Where do 403 responses cluster?
Which page class has the largest payload?
Which old PDFs continue receiving search-oriented crawler traffic?
This is where crawler identification becomes useful beyond the infrastructure team.
SEO, content, product and engineering teams can start looking at the same machine activity through different lenses.
Maintain a living crawler reference
AI crawler identities are evolving quickly.
New agents appear.
Existing providers introduce new purposes.
IP ranges change.
Verification methods evolve.
OpenAI already exposes separate agents for training, search and user-triggered requests. Anthropic does the same. Google is moving into user-triggered agents. Cloudflare’s bot classifications are increasingly purpose-based.
Your crawler reference should therefore live in a maintainable configuration or data table containing:
- provider
- stable token
- purpose
- verification method
- official source
- last-reviewed date
Do not bury it inside application logic that gets reviewed only during the next major website release.
What about bots you do not recognise?
Do not automatically classify every unusual automated user agent as an AI crawler.
Start with provider documentation.
Then use broader bot directories for discovery.
Cloudflare’sAI bot reference includes OpenAI, Anthropic and Perplexity alongside machine visitors operated by Meta, Apple, Amazon, Microsoft, ByteDance and others.
But where the provider publishes its own documentation, that should remain the stronger source for understanding what the bot actually does.
Third-party directories are useful for discovering a crawler.
The operator is the better source for its purpose.
Server logs measure access, not AI visibility
Even perfect crawler classification has a boundary.
Server logs can tell you:
- which machine requested the page
- which URL it requested
- whether the request succeeded
- how large the response was
- how long the response took
They cannot prove that:
- the page appeared in an AI answer
- it was cited
- the brand was mentioned
- the product was recommended
- a human later visited the website
Those are downstream events.
Crawler logging measures machine access.
It should not be presented as an AI visibility outcome.
That distinction is the reason the previous article, AI Bot Traffic: What AI Crawlers Reveal About Future Brand Discovery, treats bot traffic as an operational signal rather than customer demand.
Where Publive AXP Edge fits
Publive AXP Edge sits around this machine-request and delivery layer.
The value of crawler identification is not simply that another dashboard can say:
GPTBot visited 10,000 times.
Once relevant machine traffic is identified and classified, teams can inspect what those requests actually receive and identify access, rendering, latency or payload problems.
TheAXP Edge AI Crawler Optimization capability addresses that next layer: improving the representation served to relevant AI crawlers when the normal website response is incomplete or unnecessarily difficult for machines to consume.
The progression is:
identify → verify → classify → inspect → remediate
The first three tell you who the machine visitor is.
The last two tell you whether the website is actually serving it well.
The goal is not to know that GPTBot visited
At the most basic level, identifying AI crawler traffic is easy.
Search the user-agent field.
The useful version goes further.
Can you say:
which AI platforms are accessing the website?
what purpose each machine visitor serves?
whether the request is authentic?
which sections it accesses?
whether those requests succeed?
which page classes consistently fail?
That is when server logs become more than infrastructure exhaust.
They become an operational view of the machine audience.
Do not stop at “GPTBot visited us.”
Get to:
“We know which AI systems are requesting which parts of our site, why they are there, whether the requests are legitimate and whether our infrastructure is successfully serving them.”
That is when AI crawler logging becomes useful.
Frequently Asked Questions
How do I identify GPTBot in server logs?
Match the stable GPTBot token in the user-agent field, then use the IP information published inOpenAI’s crawler documentation when stronger verification is needed. GPTBot should remain separate from OAI-SearchBot and ChatGPT-User because OpenAI documents different purposes for each.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot is associated with content that may be used in model development, while OAI-SearchBot supports website discovery in ChatGPT Search. OpenAI gives them separate controls and published IP ranges.
How do I identify ClaudeBot traffic?
Match ClaudeBot, Claude-SearchBot and Claude-User separately.Anthropic’s official crawler guidance documents different purposes for the three and provides source-IP information for validation.
How do I identify PerplexityBot?
Match PerplexityBot separately from Perplexity-User, then use Perplexity’s provider documentation and published network information for verification where needed.
Can I trust an AI crawler user-agent string?
Not by itself. User agents can be spoofed. Provider-published IP ranges, reverse-DNS validation and increasingly cryptographic systems such asWeb Bot Auth provide stronger identity evidence.
Which log fields matter most?
At minimum, preserve timestamp, source IP, request path, HTTP status and user agent. Response size and response time are also useful. BothNGINX andFastly expose the underlying fields needed for this analysis.
Do crawler logs tell me whether my page was cited?
No. Logs prove that a machine requested a resource. Citation, brand mention, recommendation and human referral are separate downstream outcomes that require separate measurement.