FR, version française
Home / GEO SEO: getting cited by AI search engines / GPTBot and AI crawlers: understand and decide

GPTBot and AI crawlers: understand and decide

GPTBot is OpenAI’s web crawler. Along with ClaudeBot, PerplexityBot and the other AI providers’ crawlers, it visits your site for specific reasons. Some crawlers collect content to train a model. Others feed the index of an answer engine, and others read a page at a user’s request. Blocking the wrong crawler can remove you from ChatGPT or Perplexity answers without protecting anything. This page covers each provider’s official documentation, the useful robots.txt settings, the role of the CDN and the checks to run in server logs.

Key takeaways

GPTBot is OpenAI’s crawler that collects content to train its models. It respects robots.txt. Blocking it does not remove your site from ChatGPT search, which goes through another crawler, OAI-SearchBot. Each AI provider separates training, search index and visits requested by a user in the same way: you decide crawler by crawler.

  • Three families of AI crawlers: training, search index, retrieval on a user’s request.
  • Google-Extended is not a crawler: it is a robots.txt token, with no effect on Google Search.
  • The CDN can block a crawler before it reads robots.txt: Cloudflare changed its default settings on September 15, 2026.

What is GPTBot?

GPTBot is the OpenAI crawler that collects web pages that may be used to train its generative AI models. OpenAI describes it as a way to make its models “more helpful and safe” (OpenAI, Overview of OpenAI Crawlers). It identifies itself with GPTBot in its user-agent string and respects robots.txt rules.

A point that is often misunderstood: GPTBot does not feed ChatGPT search. That role belongs to OAI-SearchBot. OpenAI states that the two settings are independent. A site can allow OAI-SearchBot to appear in search and refuse GPTBot. If both are allowed, OpenAI may use a single crawl for both uses.

Definition

An AI crawler (or AI bot) is an automated program that visits web pages on behalf of an artificial intelligence provider. It can be used to train a model, to feed the index of an answer engine or to read a page at a user’s request.

What are the main AI crawlers and what are they used for?

Every major provider now distinguishes several crawlers. The table below is based on their official documentation, accessed on September 24, 2026. Names change often: check the provider’s page again before making any change.

ProviderTrainingSearch indexOn-demand retrieval
OpenAIGPTBotOAI-SearchBotChatGPT-User (robots.txt may not apply)
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityNo declared training crawlerPerplexityBotPerplexity-User (generally ignores robots.txt)
GoogleGoogle-Extended token (not a crawler)Googlebot, which also feeds AI Overviews and AI ModeNot relevant here
AI crawlers by family, based on each provider’s documentation (September 2026).

Anthropic: ClaudeBot and its two companions

Anthropic describes three crawlers (Anthropic, help article updated on April 7, 2026). Blocking ClaudeBot signals that your future content should be excluded from training datasets. Blocking Claude-User prevents Claude from retrieving your page when a user asks a question. Blocking Claude-SearchBot prevents your content from being indexed for its search. Anthropic also states that the rules apply subdomain by subdomain.

Perplexity: an index, no declared training

Perplexity says that PerplexityBot is used to surface and link websites in its results. It is not used to train foundation models. Perplexity-User, for its part, visits a page when a user asks a question. The provider states that it generally ignores robots.txt, since the visit is requested by a person (Perplexity, crawler documentation).

Google: Google-Extended is not a crawler

Google-Extended has no user-agent string of its own: it is a token read in robots.txt. It controls whether your content is used to train future Gemini models and for grounding in Gemini apps and the Vertex AI API. Google states that it has no impact on inclusion in Search and is not a ranking signal (Google Search Central, common crawlers). Google’s AI Overviews and AI Mode, on the other hand, rely on Googlebot.

Should you block GPTBot and training crawlers?

It is a content policy decision, to be made in writing with management. Think it through family by family.

FamilyEffect of blockingDefault recommendation
Training (GPTBot, ClaudeBot, Google-Extended)Your content is left out of future datasets. No direct measurable return, long-term effect on what the model knows about youEditorial choice: block if your content is your product (media, paid data)
Search index (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot)Your pages can no longer be cited as sources by these enginesAllow if AI visibility is one of your goals
On-demand retrieval (ChatGPT-User, Claude-User, Perplexity-User)A user can no longer have the assistant read your pageAllow: a person is requesting the page

A brand that wants to be cited by ChatGPT must at the very least let OAI-SearchBot through. Our article on ChatGPT SEO details the other levers.

How do you allow or block AI crawlers in robots.txt?

The syntax is that of any robots.txt: a group of rules for one or more user agents. Example of a “search yes, training no” policy:

# Answer engines: allowed
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# Training: refused
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
  • OpenAI indicates a delay of about 24 hours to take a robots.txt change into account;
  • each subdomain has its own robots.txt: check them all;
  • a Disallow inherited from an old policy can block a search crawler without anyone knowing.

The general rules on syntax, precedence and testing are detailed in our guide to the robots.txt file. The llms.txt file, often mentioned at this point, neither blocks nor allows any crawler: see our analysis of the llms.txt file.

Why robots.txt is not enough: the CDN layer

A CDN or a web application firewall can refuse a crawler before it reads robots.txt. This is the most common blind spot in audits. At Cloudflare, three steps have followed one another.

  1. July 1, 2025: AI crawlers blocked by default for new domains (Cloudflare press release) and launch of pay per crawl, which lets sites charge for crawling (Cloudflare, pay per crawl).
  2. September 24, 2025: content signals in robots.txt (search, ai-input, ai-train), declarative preferences with no technical enforcement (Cloudflare, Content Signals Policy).
  3. September 15, 2026: new default settings for new domains. Training crawlers and agents blocked on pages that display ads, search crawlers allowed (Cloudflare, announcement of July 1, 2026).

Warning

These default settings target new domains. They depend on the type of crawler and on whether ads are present. Never claim that “Cloudflare blocks everything by default”: open the dashboard and check the account’s actual setting.

How do you check that AI crawlers visit your site?

Reading robots.txt tells you what is allowed, not what actually happens. Three checks complete the diagnosis.

  • Server logs: filter the GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot user agents, and note the pages visited and response codes. Detailed method in our log file analysis guide.
  • IP address ranges: OpenAI publishes the ranges of each of its crawlers in JSON format. A visit claiming to be GPTBot from outside these ranges is an impostor.
  • Real access test: a request with the crawler’s user agent shows the CDN’s response (200 code or block).

Google-Extended never appears in logs, since it is not a crawler. To then track the effect of these settings on citations, see our page on how to measure AI visibility. The overall strategy is in our GEO guide.

Frequently asked questions

Does blocking GPTBot make my site disappear from ChatGPT?

No. GPTBot is used for training. ChatGPT search relies on OAI-SearchBot, which you can allow separately. OpenAI states that the two settings are independent.

What is the difference between GPTBot and ChatGPT-User?

GPTBot crawls the web automatically for training. ChatGPT-User visits a page because a user asked for it in ChatGPT or in a custom GPT. OpenAI states that robots.txt may not apply in that case.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google states that Google-Extended has no impact on inclusion in Search. AI Overviews and AI Mode rely on Googlebot. Google-Extended concerns Gemini training and grounding in Gemini and Vertex AI.

How do you know if ClaudeBot visits your site?

Filter your server logs on the ClaudeBot user agent and note the pages visited and the response codes. Also check that your CDN does not block it before robots.txt is read.

Sources

  1. OpenAI, Overview of OpenAI Crawlers, undated documentation. Accessed on September 24, 2026.
  2. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?, updated on April 7, 2026. Accessed on September 24, 2026.
  3. Perplexity, Perplexity Crawlers, undated documentation. Accessed on September 24, 2026.
  4. Google Search Central, List of Google’s common crawlers. Accessed on September 26, 2026.
  5. Cloudflare, Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large, press release of July 1, 2025. Accessed on September 24, 2026.
  6. Cloudflare, Introducing pay per crawl, published on July 1, 2025. Accessed on September 24, 2026.
  7. Cloudflare, Content Signals Policy, published on September 24, 2025. Accessed on September 24, 2026.
  8. Cloudflare, Your site, your rules: new AI traffic options for all customers, published on July 1, 2026. Accessed on September 24, 2026.

All GEO guides