Back to Insights
GEO ·27 August 2026 ·7 min read

Should you block AI crawlers or let them in?

Blocking every AI crawler is usually too blunt. GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot each serve a different purpose — how to configure robots.txt for training versus visibility.

For a business that wants to be discovered through ChatGPT, Perplexity, Gemini and other AI services, blocking every AI crawler is usually too blunt an approach. A better policy is selective: allow crawlers that support search, current answers and source citations, while deciding separately whether you want crawlers that collect content for model development or broadly reusable datasets.

That distinction matters because a single AI provider may operate several bots for completely different purposes. OpenAI, for example, separates GPTBot from OAI-SearchBot, while Anthropic distinguishes ClaudeBot, Claude-SearchBot and Claude-User. Blocking everything associated with one vendor can therefore reduce AI visibility even when your actual goal was only to opt out of model training.

AI crawlers do not all perform the same job

With traditional search engines, the robots.txt decision was relatively straightforward: should a crawler access a page or not? Generative AI has made that distinction more granular.

A training crawler collects public web content that may contribute to future models. A search crawler discovers and indexes pages so an AI product can retrieve current information and cite sources. A third category consists of user-triggered fetchers, which access a page because an individual user has initiated an action that requires it.

Those are materially different use cases. A publisher may not want its articles collected for future model training while still wanting those articles to be cited when somebody asks an AI assistant a relevant question.

Blocking GPTBot does not necessarily block ChatGPT Search

OpenAI illustrates the distinction clearly. GPTBot crawls content that may be used to improve or train generative foundation models. OAI-SearchBot serves a different purpose: it supports discovery and surfacing of websites in ChatGPT Search. OpenAI specifically advises publishers not to block OAI-SearchBot if they want their content included in summaries, snippets and links in ChatGPT search experiences.

A publisher that wants search visibility but does not want GPTBot collecting its content can therefore use:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

OpenAI also operates ChatGPT-User for certain actions initiated by ChatGPT users and Custom GPTs. Because these requests are triggered by a person rather than by automatic web crawling, normal robots.txt rules may not always apply in the same way.

The key point is simple: blocking GPTBot and blocking ChatGPT Search are separate decisions.

ClaudeBot, Claude-SearchBot and Claude-User

Anthropic now uses a comparable separation. ClaudeBot collects public web content that could contribute to model development. Claude-SearchBot navigates the web to improve search results, while Claude-User retrieves content in response to user activity in Claude. Anthropic states that blocking Claude-SearchBot may reduce a site's visibility or accuracy in Claude's search results.

A selective configuration could therefore look like this:

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

The older anthropic-ai token still appears in many crawler lists and legacy robots.txt examples. It is not included in Anthropic's current official list, which names ClaudeBot, Claude-SearchBot and Claude-User. It is therefore safer to treat anthropic-ai as a legacy compatibility rule rather than relying on it to control current Anthropic crawling.

PerplexityBot is primarily about search visibility

Perplexity presents a different trade-off. PerplexityBot is used to index websites and surface them as sources in Perplexity search results. Perplexity states that this crawler is not used to collect content for foundation-model pre-training.

For organisations pursuing AI visibility, blocking PerplexityBot therefore has a fairly direct consequence. It restricts Perplexity's ability to index the page content, but also makes it harder for the service to use that content as a current source.

Perplexity additionally operates Perplexity-User for user-initiated actions. Its documentation says this fetcher generally ignores robots.txt because the request originates from a user action rather than an autonomous crawl.

Google-Extended is not Googlebot

Google requires particular care because Google-Extended is not a conventional crawler with its own HTTP user-agent string. It is a robots.txt product token that lets publishers control whether content Google has crawled may be used for future Gemini model training and for certain grounding functions in Gemini Apps and Vertex AI.

A site can therefore use:

User-agent: Google-Extended
Disallow: /

This does not remove the site from Google Search. Google explicitly states that Google-Extended neither affects inclusion in Search nor acts as a ranking signal. Google's AI features inside Search, including AI Overviews and AI Mode, continue to rely on Googlebot and the standard Search controls.

That makes Google-Extended a useful example of why blanket AI-bot blocking is problematic: a publisher can restrict certain Gemini uses without sacrificing conventional Google Search visibility.

Where does CCBot fit?

CCBot belongs to Common Crawl, which maintains an openly accessible corpus containing raw web data, metadata and extracted text. Common Crawl currently identifies its crawler as CCBot/2.0 and states that it checks and respects robots.txt.

Blocking it is straightforward:

User-agent: CCBot
Disallow: /

The strategic consequence is less directly tied to one AI assistant. Allowing CCBot means your public content may enter a dataset that can be reused by researchers, companies and other downstream systems. Blocking it gives you more control over future collection by Common Crawl, but may also keep your content out of datasets that other systems subsequently use.

Changing robots.txt does not retroactively remove copies that have already been collected or distributed.

Robots.txt is a policy signal, not an access-control system

The syntax itself is simple. Disallow: / blocks the entire site for a matching compliant crawler, while Allow: / permits access. Rules can also target individual directories or paths. Under the Robots Exclusion Protocol, the most specific matching Allow or Disallow rule takes precedence.

However, robots.txt is not security. The file is publicly accessible, and a scraper that chooses not to comply can ignore it. Sensitive material should be protected through authentication, authorisation and, where appropriate, server or WAF controls. A user-agent header also does not by itself prove that traffic genuinely originates from the company named in that header.

It is also important to distinguish blocking crawling from blocking discovery or indexing. A service may learn that a URL exists through links or third-party indexes even when it cannot crawl the page itself. Where complete removal from an index is required, the relevant noindex and platform-specific controls should be assessed as well.

Set policy by use case, not by vendor

For many business websites, selective access is a sensible starting point. Search-oriented crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot can remain accessible where AI discovery and citations matter, while GPTBot, ClaudeBot and Google-Extended can be evaluated separately according to the organisation's position on training and reuse.

The policy does not need to be identical across an entire domain. Public knowledge-base articles may benefit from broad machine discoverability, while paid research, licensed material, private portals and other commercially sensitive sections may justify tighter controls.

Crawler policies should also be reviewed periodically. Names, functions and operator policies evolve, and new agents continue to appear. Production rules should therefore be checked against current first-party documentation rather than copied indefinitely from an old AI-bot list.

In summary

Do not treat every AI crawler as the same thing: model training, AI search and user-triggered retrieval serve different purposes. If AI visibility matters, allowing relevant search crawlers while separately evaluating training crawlers is usually more precise than blocking everything. robots.txt provides useful control over compliant automated crawlers, but it is neither a security boundary nor a retroactive deletion mechanism. The right configuration depends on what each section of your site is meant to achieve: be discoverable where visibility creates value, and restrict reuse where control matters more.

Checking your robots.txt?

Sure your robots.txt is doing the right thing?

Neuralex runs a site through 15+ GEO checks — from crawler access to how citeable the text actually is — and delivers a score with concrete fixes, including a robots.txt that opts out of training without costing you AI visibility.