Guide

AI crawlers: who they are, what they do and how to control them

AI companies send separate crawlers to train models, to build AI search indexes and to fetch pages while answering someone. Each has its own robots.txt name, so you can choose which ones read your site.

Last checked against official documentation on October 4, 2026.

Three jobs, three kinds of crawler

AI companies don't send one crawler. They send different agents for different jobs, and blocking one doesn't block the others.

JobWhat the crawler doesExamples
TrainingCollects pages that may be used to train future modelsGPTBot, ClaudeBot, Meta-ExternalAgent, MistralAI-Training
AI searchBuilds the index that AI search and answers draw onOAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer
AnswersFetches a page while an assistant answers someone, because their request needs itChatGPT-User, Claude-User, Perplexity-User, Google-Agent, DuckAssistBot

Google works differently. AI Overviews and AI Mode are part of Google Search, so robots.txt rules for Googlebot control how they crawl, and since August 31, 2026, a Search Console setting can opt a site out of them. Google-Extended isn't a crawler at all, but a robots.txt name that controls whether pages Google already crawls may be used to train Gemini models and to ground Gemini's answers. The Google-Extended guide covers both.

The main AI crawlers

CompanyAgentJobFollows robots.txtIP list
OpenAIGPTBotTrainingYesgptbot.json
OpenAIOAI-SearchBotChatGPT searchYessearchbot.json
OpenAIChatGPT-UserAnswersMay not apply, says OpenAIchatgpt-user.json
AnthropicClaudeBotTrainingYesbots.json
AnthropicClaude-SearchBotClaude searchYesbots.json
AnthropicClaude-UserAnswersYesbots.json
PerplexityPerplexityBotPerplexity search, not model trainingYesperplexitybot.json
PerplexityPerplexity-UserAnswersGenerally not, says Perplexityperplexity-user.json
GoogleGoogle-ExtendedGemini training and grounding (a robots.txt name, not a crawler)YesUses Google's existing crawlers
GoogleGoogle-AgentAgents acting on a person's requestGenerally not, says Googleuser-triggered-agents.json
AmazonAmazonbotImproving Amazon's products; may train Amazon AI modelsYes, but not Crawl-delayAmazonbot list
AmazonAmzn-SearchBotSearch experiences such as Alexa, not trainingYes, but not Crawl-delayAmzn-SearchBot list
AmazonAmzn-UserAnswers, such as Alexa questionsMay not follow every directive, says AmazonAmzn-User list
AppleApplebotSearch in Siri, Spotlight and Safari; may train Apple modelsYes, but not Crawl-delayapplebot.json
AppleApplebot-ExtendedApple model training (a robots.txt name, not a crawler)YesUses Applebot's addresses
MetaMeta-ExternalAgentUses such as AI model training and direct indexingYesMeta's AS32934, via whois
MetaMeta-WebIndexerMeta AI search and citationsYesMeta's AS32934, via whois
MetaMeta-ExternalFetcherFetches for users and AI agent tasksMay bypass it, says MetaMeta's AS32934, via whois
ByteDanceBytespiderNot documented by ByteDanceNot reliably, say independent reportsNone published
Common CrawlCCBotOpen web archive, widely used to train AI modelsYes, including Crawl-delayccbot.json
DuckDuckGoDuckAssistBotAnswers in DuckAssist, not trainingYes, within 72 hoursOn its help page
MistralMistralAI-UserAnswers in Vibe, formerly Le ChatMistral says it governs which sites it fetchesmistralai-user-ips.json
MistralMistralAI-IndexMistral search, not trainingNot statedmistralai-index-ips.json
MistralMistralAI-TrainingTraining Mistral modelsYesNone published

Perplexity adds one detail: when robots.txt blocks PerplexityBot, Perplexity may still index the page's domain, headline and a brief factual summary.

Guides to the most searched-for ones:

To see which of these crawlers your robots.txt lets read a page, and the line that decides it, use the free robots.txt checker for AI crawlers. To write a robots.txt that allows or blocks each of them, use the free robots.txt generator, and to draft an llms.txt from your own pages, the free llms.txt generator.

Other AI agents you may see

  • Diffbot crawls for Diffbot's Knowledge Graph and web search, which Diffbot says isn't used for AI training. It follows robots.txt, including Crawl-delay, by default, but Diffbot's customers can switch that off for crawls they run with its software and can send their own user agent.
  • KimiBot, Kimi-SearchBot and Kimi-User (Moonshot AI) collect training data, build Kimi's search and fetch pages for users. Moonshot publishes an IP list for each, and says robots.txt rules may not directly apply to Kimi-User.
  • AI2Bot (Allen Institute for AI) collects web content used to train open language models.
  • YouBot (You.com) is no longer documented by You.com: about.you.com/youbot, the page crawler directories cite for it, returns a 404 error, so requests that use the name can't be verified.
  • xAI's Grok and DeepSeek publish no crawler documentation we could find.

Opt out of training without leaving AI answers

These rules ask OpenAI, Anthropic and Google not to use your pages for model training, and leave ChatGPT search, Claude's search and the answers that fetch your pages alone:

TEXT
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: MistralAI-Training
Disallow: /

Two things to know before you copy them:

  • Google-Extended does more than training. Google says it also controls grounding, the content Gemini Apps and Grounding with Google Search on Vertex AI pull from Google's index while answering. Blocking it doesn't affect Google Search, AI Overviews or AI Mode.
  • Perplexity has no training crawler to block. Perplexity says it doesn't build foundation models, so your content isn't used for pre-training. PerplexityBot only feeds Perplexity's search.
  • Some crawlers do more than training. Amazonbot, Meta-ExternalAgent and CCBot collect data that may be used for training, but each serves other purposes too, so read their guides before you block them.

A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them. The free robots.txt generator writes these groups for you and repeats your private paths in each group that lets a crawler in.

Agents that act for a person

The agents that fetch a page because someone asked for it often don't follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says its user-triggered fetchers, including Google-Agent, generally ignore them, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic is the exception: it says its bots honor robots.txt, and that blocking Claude-User stops Claude from retrieving your pages for people's questions.

Keeping these agents out takes a firewall rule rather than robots.txt, and it also keeps your pages out of the answers people ask for. Think about that trade before you block them.

How to verify an AI crawler

User agent strings are easy to fake, so a request that says GPTBot or ClaudeBot proves nothing on its own. Every company in the table above publishes the IP addresses its agents use. Check the request's source IP against the right list: when it matches, the request is real. For Google's common crawlers, Google also documents a reverse DNS check: their hostnames match crawl-***-***-***-***.googlebot.com.

Where to see AI crawlers

Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up. Search them for the agent names in the table above.

Make sure the content you want read is in the HTML your server sends. Google says it can process content in JavaScript as long as it isn't blocked, but OpenAI, Anthropic and Perplexity don't say whether their crawlers run JavaScript, and a page that only fills in after scripts run may look empty to them.

OneLence AI visibility turns those requests into a report when your server reports them to OneLence. It matches each request to the crawler that sent it, marks the ones confirmed by the company's IP list as verified, and shows which pages each crawler read, which pages assistants fetched while answering people, and the visitors those assistants sent you and their conversions.

Frequently asked questions

What is an AI crawler?

An AI crawler is an automated client that an AI company sends to websites. Some collect pages to train models (GPTBot, ClaudeBot), some build the index behind AI search (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and some fetch a page while an assistant answers someone (ChatGPT-User, Claude-User, Perplexity-User).

Will blocking AI training crawlers remove my site from ChatGPT or Claude?

No. GPTBot and ClaudeBot only collect training data. ChatGPT search relies on OAI-SearchBot and Claude's search on Claude-SearchBot, and each has its own robots.txt name, so a rule for a training crawler doesn't apply to them. Google-Extended is the exception to keep in mind: besides Gemini training, it also controls grounding in Gemini Apps and Vertex AI.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google says Google-Extended doesn't affect inclusion or ranking in Google Search, and AI Overviews and AI Mode are part of Search. To opt out of them, use the Search generative AI control in Search Console, which Google rolled out to all websites worldwide as of August 31, 2026. It removes your links and content from AI Overviews, AI Mode and generative AI features in Discover without affecting the rest of Search.

Which AI crawlers ignore robots.txt?

Mostly the ones that act for a person. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says the same of its user-triggered fetchers such as Google-Agent, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic says its bots, Claude-User included, honor robots.txt. Bytespider is the exception among the automatic crawlers: several independent reports say it doesn't reliably follow robots.txt.

How do I know a request really comes from an AI company?

Check the source IP address against the list the company publishes, such as openai.com/gptbot.json, claude.com/crawling/bots.json or perplexity.com/perplexitybot.json. Anyone can copy a crawler's user agent string, so the name alone proves nothing.

Why don't AI crawlers show up in Google Analytics?

Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up.