> Bron: https://neuralex.nl/en/blog/ai-crawlers-herkennen
> GPTBot, ClaudeBot, PerplexityBot and more do not all do the same thing: some gather training data, some build a search index, some fetch a page live at a user’s request. How to filter and verify them in your logs.

[Back to Insights](/en/blog)

GEO ·29 August 2026 ·7 min read

# Which AI crawlers visit your site — and how do you spot them in your logs?

GPTBot, ClaudeBot, PerplexityBot and more do not all do the same thing: some gather training data, some build a search index, some fetch a page live at a user’s request. How to filter and verify them in your logs.

The main AI-related bots you may see in server or CDN logs include OpenAI's GPTBot and OAI-SearchBot, Anthropic's ClaudeBot and Claude-User, Googlebot alongside Google-Extended, PerplexityBot and Perplexity-User, Bytespider, Amazonbot, CCBot and Applebot. Some collect content for possible model development, some maintain search indexes, and some fetch a page because a user has just asked an AI system a question.

That distinction should drive your policy. You may block a training crawler while allowing a search or citation crawler to keep public pages discoverable — see also [should you block AI crawlers or let them in](/en/blog/block-ai-crawlers-or-let-them-in). The practical approach is to combine robots.txt policy with real log monitoring.

## Training crawlers, search crawlers and live fetchers

It is useful to think in three categories. Training-oriented crawlers gather public web content that may feed future datasets or model improvement. Search crawlers build or refresh an index used to discover relevant pages later. User-triggered fetchers retrieve a page in response to an individual action or question.

OpenAI documents this separation clearly. GPTBot crawls content that may be used to train generative AI foundation models. OAI-SearchBot exists for ChatGPT search and is used to surface sites in search answers. OpenAI also uses ChatGPT-User for certain user-triggered actions; that agent is not an automatic web crawler. The agents have separate user-agent strings and published IP ranges.

Anthropic follows a similar model. ClaudeBot can collect public web content for model development, while Claude-User is used when Claude accesses a site at a user's direction. Anthropic also documents Claude-SearchBot for search-related discovery. A request containing "Claude" therefore does not automatically tell you whether it is training, search, or live retrieval traffic.

Perplexity separates PerplexityBot from Perplexity-User. PerplexityBot indexes pages so they can be surfaced and linked in Perplexity search results and, according to Perplexity, is not used to crawl content for foundation-model training. Perplexity-User may fetch a page during a user query and include a link to that page in the answer.

## Google-Extended and Applebot-Extended are control signals

Google-Extended often appears in "AI crawler" lists, but it is not a separate HTTP crawler. Google states that Google-Extended has no independent request user-agent string. Existing Google crawlers perform the crawling, while the Google-Extended robots.txt token controls whether crawled content can be used for future Gemini model training and certain grounding use cases. Blocking Google-Extended does not remove a site from Google Search.

Applebot-Extended is similar. Applebot is the crawler that fetches pages. Applebot-Extended itself does not crawl webpages; it is a robots.txt control that lets publishers restrict use of Applebot-collected content for training Apple's generative foundation models. Searching access logs only for "Applebot-Extended" can therefore produce a misleading zero even while Applebot is active.

## Other AI-related crawlers worth tracking

Amazon says Amazonbot is used to improve its products and services and that crawled material may be used to train Amazon AI models. Amazon now also distinguishes Amzn-SearchBot, used for search experiences, from Amzn-User, which can fetch current information in response to a user action.

CCBot belongs to Common Crawl rather than one AI assistant. It builds an open web dataset reusable for search, research and machine learning, making it indirectly relevant to many AI systems. Common Crawl publishes IP ranges and reverse-DNS information for verification.

Bytespider is widely identified as a ByteDance crawler and is classified by Cloudflare as an AI crawler. The caveat is transparency: ByteDance does not provide the same level of first-party crawler documentation, purpose statements and verification material as some other providers. For operational purposes, matching the Bytespider token is useful, but claims about its exact role should be made cautiously unless you have stronger evidence.

## How to find AI crawlers in server logs

Apache and Nginx access logs commonly record the HTTP user-agent with each request. A basic first pass is therefore a case-insensitive filter for known tokens such as GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|PerplexityBot|Perplexity-User|Bytespider|Amazonbot|CCBot|Applebot. With grep, that is enough to isolate likely AI-related requests from a conventional access log.

For counts by user-agent, awk can extract the user-agent field from a standard combined log format before sorting and counting it. Check your own format first because proxies and custom configurations may place fields differently. Do not stop at request counts. Record source IP, timestamp, path, status, bytes transferred and crawl pattern. A crawler fetching sitemaps and articles behaves differently from a bot probing admin endpoints or private-looking paths.

## A user-agent string is not proof

Any HTTP client can claim to be Googlebot, GPTBot or almost anything else. User-agent matching is therefore useful for discovery, but weak as an authentication mechanism.

Google recommends verifying crawler traffic through reverse DNS and then performing a forward lookup to confirm that the hostname resolves back to the original IP. Google also publishes crawler IP ranges. OpenAI, Perplexity, Amazon, Common Crawl and Apple publish verification information or address ranges for relevant crawlers as well. This matters most when you create WAF allow rules. An unconditional allow rule based only on a familiar user-agent can let a spoofing bot bypass controls intended for untrusted automation.

## Monitoring at scale

For a small site, recurring log filters may be sufficient. Larger environments benefit from dashboards that group activity by provider, path, response code and time period. Cloudflare's AI Crawl Control can break AI-related traffic down by crawler, operator, hostname, path, status code and transferred data. Cloudflare also notes that user-agent-based detection can be spoofed, which is why stronger bot-identification signals are preferable when available.

## Why businesses should care

Crawler logs are not a complete measure of AI visibility, but they are useful technical evidence. If OAI-SearchBot, PerplexityBot or comparable search agents never receive successful responses from important pages, it is worth checking robots.txt, your CDN, WAF and origin configuration. The opposite is also true: heavy crawling does not prove that an AI assistant frequently cites your brand.

The key is to avoid treating all AI bots as one category. A company may have copyright, privacy or infrastructure reasons to restrict training-oriented crawlers while still allowing search and user-triggered agents to access public knowledge-base content. That can preserve discoverability without granting every crawler the same access. Logs also help expose impersonators. A bot claiming to be a known AI crawler but arriving from inconsistent infrastructure, ignoring expected behavior or aggressively scanning sensitive paths should not be trusted merely because its user-agent contains the name of a well-known AI company.

## Conclusion

The useful question is not simply which AI crawlers visit your site, but what each one is trying to do and whether the request is genuine. Filter logs for known user-agent tokens, verify important traffic with DNS or published IP ranges where possible, and separate training, search indexing and user-triggered retrieval. That gives you a defensible basis for deciding which crawlers to block, which to allow and which deserve closer monitoring.

Measuring visibility in AI answers

## Want to know which AI crawlers are actually visiting your site?

We set up log monitoring and verification so you can decide, with evidence, who to allow and who to block. Curious how that would look for your site?

[Ask your question](/en/contact) [Read should you block AI crawlers](/en/blog/block-ai-crawlers-or-let-them-in)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/ai-crawlers-herkennen
