BlogAI search

AI crawlers explained: GPTBot, ClaudeBot, PerplexityBot and robots.txt

Benjamin Libor9 min read

AI crawlers come in three kinds, and the difference decides whether blocking one costs you visibility. Training crawlers such as OpenAI's `GPTBot` collect pages to train models; search crawlers such as `OAI-SearchBot` and `PerplexityBot` build the index an assistant searches when it answers; and fetchers such as `ChatGPT-User` read a page live when a user's question needs it. Blocking a search crawler removes your pages from the answers. Blocking a training crawler does not.

In short

  • There are three kinds of AI bot: training crawlers, search-index crawlers and live fetchers, with different names and different consequences.
  • Blocking the search crawler (OAI-SearchBot, PerplexityBot, Googlebot) takes you out of that engine's answers. Blocking the training crawler (GPTBot, Google-Extended, CCBot) does not.
  • A wildcard Disallow: / or a bot-protection rule that challenges crawlers is the most common way brands disappear from AI answers by accident.
  • Server logs, a CDN's analytics or a server-side tag show which crawlers visit, which pages they fetch and which they never reach.
  • Content that only appears after JavaScript runs is invisible to most crawlers, whatever robots.txt says.

Three kinds of AI bot

Every AI company that fetches web pages does it for one of three reasons, and most document a separate bot for each.

Training crawlers collect pages to train a model. What they collect may show up, blurred, in what the model "knows" months later. Blocking them is a policy decision about training; it has little effect on whether an assistant can cite you today.

Search-index crawlers build the index an assistant queries when it searches the web to answer. This is the one that matters for visibility. If your pages are not in the index, they cannot be read, quoted or linked.

Fetchers read a specific page at the moment a user's question needs it: when a user pastes a link, or when the assistant wants the current version of a page it found in search. Several companies state that fetchers do not follow robots.txt, because they act on a person's request in the same way a browser does.

A simple test: if a bot exists so the company can answer a question right now, blocking it hurts your visibility. If it exists so the company can train a model later, blocking it is a separate choice.

The bots the companies document

The names below are published in each company's own documentation at the time of writing. Names and roles change, so check before you edit robots.txt.

CompanyBotKindBlocking it means
OpenAIGPTBotTrainingYour pages are not used to train OpenAI's models. ChatGPT answers are not affected.
OpenAIOAI-SearchBotSearch indexYour pages drop out of ChatGPT search, so ChatGPT cannot find or cite them.
OpenAIChatGPT-UserFetcherA pasted link or a found result cannot be read live.
AnthropicClaudeBotCrawlerAnthropic's documented crawler. It has since added separate agents for user fetches and search; check its documentation for current names.
PerplexityPerplexityBotSearch indexYour pages drop out of Perplexity's index and its answers.
PerplexityPerplexity-UserFetcherPerplexity states this agent generally ignores robots.txt because it acts on a user's request.
GoogleGooglebotSearch indexYou leave Google Search, and with it AI Overviews and AI Mode.
GoogleGoogle-ExtendedTraining controlA token, not a crawler. Google says it controls use for Gemini models and grounding and does not affect Search or AI Overviews.
Common CrawlCCBotOpen datasetYour pages leave the Common Crawl corpus, which many models train on. No direct effect on answers.
MetaMeta-ExternalAgent and othersMixedAgents for training and for its assistant's search; names have changed, so check Meta's documentation.
AppleApplebot, Applebot-ExtendedSearch and training controlApplebot feeds Siri and Spotlight; Applebot-Extended controls training use, like Google-Extended.

Three details are easy to miss:

  • Google-Extended and Applebot-Extended are controls, not crawlers. Googlebot still fetches the page; the token only says what Google may do with it.
  • Fetchers mostly ignore robots.txt by design. A Disallow line will not stop a user from pasting your URL into ChatGPT or Perplexity.
  • Names can be faked. Anyone can set a user-agent string to GPTBot. OpenAI, Google, Anthropic and Perplexity publish the IP ranges their bots use, so a serious check verifies the address.

What each one means for visibility

The question marketers actually ask is "if I block this, do I lose mentions?". By engine:

  • ChatGPT: mentions from training memory depend on what the model learned, which you cannot change quickly. Mentions with citations depend on OAI-SearchBot finding your pages. Blocking GPTBot changes neither in the short term. See how ChatGPT decides which pages to cite.
  • Perplexity: almost every answer comes from a live search, so blocking PerplexityBot is close to opting out. Our Perplexity playbook covers the rest.
  • Google AI Overviews and AI Mode: they draw on the Google Search index. If Googlebot can index a page, it is eligible; Google-Extended does not change that. See how to show up in Google's AI Overviews.

How to decide what to allow

Most tech companies want the same thing: be findable in every assistant, keep the option to opt out of training, and keep private areas private.

You wantAllowBlock or leave to policy
To be cited in every assistantOAI-SearchBot, PerplexityBot, Googlebot, ClaudeBot and the other search crawlersNothing that answers questions
To opt out of model training but stay visibleThe search crawlers aboveGPTBot, CCBot, Google-Extended, Applebot-Extended, Meta's training agent
To allow everythingAll of themNothing
To keep some areas privateThe search crawlers, on public pages onlyEverything on /app/, /account/, staging domains and internal docs

Two notes. A robots.txt line is a request, not a lock; a page that must not be read needs authentication. And opting out of training today does not remove content collected earlier. For a marketing site and documentation, the simplest defensible position is: allow all search crawlers on public pages, block training crawlers if your policy says so, and block every bot from the app.

Example robots.txt lines

robots.txt sits at the root of the site and is read in blocks; each rule applies to the User-agent lines directly above it. Here is a complete, conservative file for a made-up website builder, Allsite.

The first block allows the search crawlers on public pages and keeps the app private:

  • User-agent: OAI-SearchBot
  • User-agent: PerplexityBot
  • User-agent: ClaudeBot
  • User-agent: Googlebot
  • Allow: /
  • Disallow: /app/
  • Disallow: /account/

The second block opts out of training, which is a policy decision:

  • User-agent: GPTBot
  • User-agent: CCBot
  • User-agent: Google-Extended
  • User-agent: Applebot-Extended
  • Disallow: /

The third block covers everyone else, and the last line points every crawler at the sitemap:

  • User-agent: *
  • Allow: /
  • Disallow: /app/
  • Disallow: /account/
  • Sitemap: https://www.allsite.example/sitemap.xml

To allow training too, delete the second block; to opt out of one company only, keep just its bot there. Whatever you choose, keep the Sitemap: line. After editing, open https://yourdomain/robots.txt in a browser to confirm the live file matches: many CMSs and CDNs generate or rewrite the file, and the version on disk is not always the version crawlers see.

How to see which crawlers actually visit

robots.txt says what you permit. Whether the bots come, how often, and which pages they read is answered by your traffic data, in one of three places.

Server logs

Every request is recorded with its user-agent string, the page, the response code and the time. Filter for the bot names above and you have the complete picture. The catch is access: hosted platforms (Webflow, Framer, managed WordPress) may not give you raw logs at all.

Your CDN's analytics

If the site sits behind Cloudflare, Fastly, Vercel or a similar network, their dashboards show traffic by user agent, and some have a dedicated AI crawler view with per-bot counts. This is the quickest route when you have it.

A tag on the page

An ordinary analytics tag records only clients that run JavaScript, which most crawlers do not, so it undercounts bots badly. A server-side or edge tag that records each request before the page is served solves this and works on hosted platforms without logs. In Echo this is the AI crawlers page, fed by a snippet, an uploaded log file or Echo's own edge script.

What to look for

Whichever source you use, ask:

  1. Which bots come at all? A site OAI-SearchBot has never visited has a crawling problem before it has a content problem.
  2. Which pages do they fetch, and which never? Compare the list against the pages you want cited. Crawlers follow links; a page with no internal links and no sitemap entry can be missed for months.
  3. What response codes do they get? A run of 403 or 429 responses, or a 200 that is actually a challenge page, means the crawler is turned away while your robots.txt says welcome.

A worked example: Fernwick, a made-up scheduling tool, saw ChatGPT citing a rival's "Fernwick vs" page but never Fernwick's own comparison page. The logs showed OAI-SearchBot fetched the homepage and pricing page weekly and the comparison page never; it was linked only from a blog footer. One link from the pricing page and a sitemap entry later, the crawler found it within days.

The mistakes that remove you from AI answers

  • A wildcard that blocks everything. A block meant for one bot, written as User-agent: * followed by Disallow: /, silently blocks every crawler including Google. Check the whole file, not just the lines you added.
  • Blocking the search crawler when you meant the training one. Opting out of training with a block that also lists OAI-SearchBot removes you from ChatGPT search.
  • A bot-protection rule that challenges crawlers. Firewalls and bot managers often ship with rules that serve a JavaScript challenge or a 403 to anything automated, and known AI crawlers get caught. Most products have an allowlist for verified bots; put the search crawlers on it.
  • A staging rule that went live. Staging sites block everything, and the rule sometimes ships at launch.
  • Content behind JavaScript. If the page is an empty shell until a script runs, a crawler that does not execute scripts sees nothing. View the page source: the text you want cited should be in it.
  • Pages the crawler cannot reach. No internal links, no sitemap entry, or a sitemap that lists only the blog.
  • Slow or failing pages. A page that responds slowly or returns errors gets fetched less, then not at all.

Fixing these takes an hour for most sites, and is the first thing to check when AI visibility is lower than your rankings would suggest.

Questions people ask

Should I block GPTBot?

That is a policy decision about training, not visibility. OpenAI's documentation says GPTBot is used for training and that ChatGPT search uses OAI-SearchBot. Keep OAI-SearchBot allowed if you want to be cited; block or allow GPTBot according to your view on training.

Does blocking Google-Extended remove me from AI Overviews?

No. Google's documentation describes Google-Extended as a control over whether your content is used for Gemini models and grounding, and says it does not affect inclusion in Google Search, including AI Overviews. AI Overviews depend on Googlebot and ordinary indexing.

How do I know if an AI crawler is really from that company?

Check the IP address, not just the name. OpenAI, Google, Anthropic and Perplexity publish the IP ranges their crawlers use; a request claiming to be GPTBot from outside those ranges is something else wearing the name.

Can I see AI crawler visits in Google Analytics?

No. GA4 records browsers that run its tag, and crawlers do not run it. Use server logs, your CDN's analytics or a server-side tag. GA4 answers a different question, people who arrive from an AI answer, covered in how to see traffic from ChatGPT, Perplexity and Gemini in GA4.

Written by Benjamin Libor, founder, echo.

Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.

See what AI says about your brand. Free, in about a minute.

The free AI score checks how ready your website is for AI assistants. Echo itself tracks your customers' questions across the assistants every day and turns the answers into actions.