Three jobs, three kinds of crawler
AI companies don't send one crawler. They send different agents for different jobs, and blocking one doesn't block the others.
| Job | What the crawler does | Examples |
|---|---|---|
| Training | Collects pages that may be used to train future models | GPTBot, ClaudeBot, Meta-ExternalAgent, MistralAI-Training |
| AI search | Builds the index that AI search and answers draw on | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer |
| Answers | Fetches a page while an assistant answers someone, because their request needs it | ChatGPT-User, Claude-User, Perplexity-User, Google-Agent, DuckAssistBot |
Google works differently. AI Overviews and AI Mode are part of Google Search, so robots.txt rules for Googlebot control how they crawl, and since August 31, 2026, a Search Console setting can opt a site out of them. Google-Extended isn't a crawler at all, but a robots.txt name that controls whether pages Google already crawls may be used to train Gemini models and to ground Gemini's answers. The Google-Extended guide covers both.
The main AI crawlers
| Company | Agent | Job | Follows robots.txt | IP list |
|---|---|---|---|---|
| OpenAI | GPTBot | Training | Yes | gptbot.json |
| OpenAI | OAI-SearchBot | ChatGPT search | Yes | searchbot.json |
| OpenAI | ChatGPT-User | Answers | May not apply, says OpenAI | chatgpt-user.json |
| Anthropic | ClaudeBot | Training | Yes | bots.json |
| Anthropic | Claude-SearchBot | Claude search | Yes | bots.json |
| Anthropic | Claude-User | Answers | Yes | bots.json |
| Perplexity | PerplexityBot | Perplexity search, not model training | Yes | perplexitybot.json |
| Perplexity | Perplexity-User | Answers | Generally not, says Perplexity | perplexity-user.json |
Google-Extended | Gemini training and grounding (a robots.txt name, not a crawler) | Yes | Uses Google's existing crawlers | |
Google-Agent | Agents acting on a person's request | Generally not, says Google | user-triggered-agents.json | |
| Amazon | Amazonbot | Improving Amazon's products; may train Amazon AI models | Yes, but not Crawl-delay | Amazonbot list |
| Amazon | Amzn-SearchBot | Search experiences such as Alexa, not training | Yes, but not Crawl-delay | Amzn-SearchBot list |
| Amazon | Amzn-User | Answers, such as Alexa questions | May not follow every directive, says Amazon | Amzn-User list |
| Apple | Applebot | Search in Siri, Spotlight and Safari; may train Apple models | Yes, but not Crawl-delay | applebot.json |
| Apple | Applebot-Extended | Apple model training (a robots.txt name, not a crawler) | Yes | Uses Applebot's addresses |
| Meta | Meta-ExternalAgent | Uses such as AI model training and direct indexing | Yes | Meta's AS32934, via whois |
| Meta | Meta-WebIndexer | Meta AI search and citations | Yes | Meta's AS32934, via whois |
| Meta | Meta-ExternalFetcher | Fetches for users and AI agent tasks | May bypass it, says Meta | Meta's AS32934, via whois |
| ByteDance | Bytespider | Not documented by ByteDance | Not reliably, say independent reports | None published |
| Common Crawl | CCBot | Open web archive, widely used to train AI models | Yes, including Crawl-delay | ccbot.json |
| DuckDuckGo | DuckAssistBot | Answers in DuckAssist, not training | Yes, within 72 hours | On its help page |
| Mistral | MistralAI-User | Answers in Vibe, formerly Le Chat | Mistral says it governs which sites it fetches | mistralai-user-ips.json |
| Mistral | MistralAI-Index | Mistral search, not training | Not stated | mistralai-index-ips.json |
| Mistral | MistralAI-Training | Training Mistral models | Yes | None published |
Perplexity adds one detail: when robots.txt blocks PerplexityBot, Perplexity may still index the page's domain, headline and a brief factual summary.
Guides to the most searched-for ones:
- ClaudeBot: Anthropic's training crawler, and how it differs from Claude-User and Claude-SearchBot.
- ChatGPT-User vs GPTBot vs OAI-SearchBot: OpenAI's three agents, and which one decides whether you appear in ChatGPT search.
- PerplexityBot and Perplexity-User: Perplexity's search crawler and the agent that fetches pages for answers.
- Google-Extended: what it controls, and the Search Console setting for AI Overviews and AI Mode.
- Amazonbot: Amazon's three agents, and how to opt out of training but stay in Alexa.
- Applebot and Applebot-Extended: Apple's crawler and its training opt-out.
- Meta-ExternalAgent: Meta's five crawlers, and how to block training without breaking link previews.
- Bytespider: ByteDance's crawler, and why robots.txt may not stop it.
- CCBot: Common Crawl's crawler, and what blocking it changes.
- llms.txt: the proposed file for AI agents, and who actually reads it.
To see which of these crawlers your robots.txt lets read a page, and the line that decides it, use the free robots.txt checker for AI crawlers. To write a robots.txt that allows or blocks each of them, use the free robots.txt generator, and to draft an llms.txt from your own pages, the free llms.txt generator.
Other AI agents you may see
- Diffbot crawls for Diffbot's Knowledge Graph and web search, which Diffbot says isn't used for AI training. It follows robots.txt, including Crawl-delay, by default, but Diffbot's customers can switch that off for crawls they run with its software and can send their own user agent.
- KimiBot, Kimi-SearchBot and Kimi-User (Moonshot AI) collect training data, build Kimi's search and fetch pages for users. Moonshot publishes an IP list for each, and says robots.txt rules may not directly apply to Kimi-User.
- AI2Bot (Allen Institute for AI) collects web content used to train open language models.
- YouBot (You.com) is no longer documented by You.com: about.you.com/youbot, the page crawler directories cite for it, returns a 404 error, so requests that use the name can't be verified.
- xAI's Grok and DeepSeek publish no crawler documentation we could find.
Opt out of training without leaving AI answers
These rules ask OpenAI, Anthropic and Google not to use your pages for model training, and leave ChatGPT search, Claude's search and the answers that fetch your pages alone:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: MistralAI-Training
Disallow: /
Two things to know before you copy them:
- Google-Extended does more than training. Google says it also controls grounding, the content Gemini Apps and Grounding with Google Search on Vertex AI pull from Google's index while answering. Blocking it doesn't affect Google Search, AI Overviews or AI Mode.
- Perplexity has no training crawler to block. Perplexity says it doesn't build foundation models, so your content isn't used for pre-training. PerplexityBot only feeds Perplexity's search.
- Some crawlers do more than training. Amazonbot, Meta-ExternalAgent and CCBot collect data that may be used for training, but each serves other purposes too, so read their guides before you block them.
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them. The free robots.txt generator writes these groups for you and repeats your private paths in each group that lets a crawler in.
Agents that act for a person
The agents that fetch a page because someone asked for it often don't follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says its user-triggered fetchers, including Google-Agent, generally ignore them, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic is the exception: it says its bots honor robots.txt, and that blocking Claude-User stops Claude from retrieving your pages for people's questions.
Keeping these agents out takes a firewall rule rather than robots.txt, and it also keeps your pages out of the answers people ask for. Think about that trade before you block them.
How to verify an AI crawler
User agent strings are easy to fake, so a request that says GPTBot or ClaudeBot proves nothing on its own. Every company in the table above publishes the IP addresses its agents use. Check the request's source IP against the right list: when it matches, the request is real. For Google's common crawlers, Google also documents a reverse DNS check: their hostnames match crawl-***-***-***-***.googlebot.com.
Where to see AI crawlers
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up. Search them for the agent names in the table above.
Make sure the content you want read is in the HTML your server sends. Google says it can process content in JavaScript as long as it isn't blocked, but OpenAI, Anthropic and Perplexity don't say whether their crawlers run JavaScript, and a page that only fills in after scripts run may look empty to them.
OneLence AI visibility turns those requests into a report when your server reports them to OneLence. It matches each request to the crawler that sent it, marks the ones confirmed by the company's IP list as verified, and shows which pages each crawler read, which pages assistants fetched while answering people, and the visitors those assistants sent you and their conversions.
Frequently asked questions
What is an AI crawler?
An AI crawler is an automated client that an AI company sends to websites. Some collect pages to train models (GPTBot, ClaudeBot), some build the index behind AI search (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and some fetch a page while an assistant answers someone (ChatGPT-User, Claude-User, Perplexity-User).
Will blocking AI training crawlers remove my site from ChatGPT or Claude?
No. GPTBot and ClaudeBot only collect training data. ChatGPT search relies on OAI-SearchBot and Claude's search on Claude-SearchBot, and each has its own robots.txt name, so a rule for a training crawler doesn't apply to them. Google-Extended is the exception to keep in mind: besides Gemini training, it also controls grounding in Gemini Apps and Vertex AI.
Does blocking Google-Extended remove my site from AI Overviews?
No. Google says Google-Extended doesn't affect inclusion or ranking in Google Search, and AI Overviews and AI Mode are part of Search. To opt out of them, use the Search generative AI control in Search Console, which Google rolled out to all websites worldwide as of August 31, 2026. It removes your links and content from AI Overviews, AI Mode and generative AI features in Discover without affecting the rest of Search.
Which AI crawlers ignore robots.txt?
Mostly the ones that act for a person. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says the same of its user-triggered fetchers such as Google-Agent, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic says its bots, Claude-User included, honor robots.txt. Bytespider is the exception among the automatic crawlers: several independent reports say it doesn't reliably follow robots.txt.
How do I know a request really comes from an AI company?
Check the source IP address against the list the company publishes, such as openai.com/gptbot.json, claude.com/crawling/bots.json or perplexity.com/perplexitybot.json. Anyone can copy a crawler's user agent string, so the name alone proves nothing.
Why don't AI crawlers show up in Google Analytics?
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up.
