ClaudeBot is one of three web agents Anthropic runs. Each has its own job and its own robots.txt token, so you can allow one and block another.
Anthropic's three bots
| Bot | What it does | What blocking it in robots.txt changes |
|---|---|---|
ClaudeBot | Collects public web content that could contribute to training Claude models | Tells Anthropic to exclude your pages from future training data |
Claude-User | Visits a page when someone asks Claude a question that needs it | Claude can't retrieve your pages to answer people's questions |
Claude-SearchBot | Crawls the web to improve the quality of Claude's search results | Your pages can be less visible and less accurate in Claude's search results |
Anthropic publishes the IP addresses its bots use in one list, claude.com/crawling/bots.json.
Does ClaudeBot follow robots.txt?
Yes. Anthropic says its bots honor industry standard robots.txt directives, support Crawl-delay, and don't try to get around CAPTCHAs or other anti-bot measures. robots.txt works per host, so add the rules to the robots.txt of each subdomain you want covered.
Anthropic advises against blocking its bots by IP address. It may not work as intended, and it can stop the bots from reading your robots.txt at all.
robots.txt rules for ClaudeBot
Block ClaudeBot from the whole site:
User-agent: ClaudeBot
Disallow: /
Slow it down instead of blocking it:
User-agent: ClaudeBot
Crawl-delay: 1
Opt out of training while keeping Claude's answers and search:
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
How to verify a ClaudeBot request
The user agent of a ClaudeBot request contains ClaudeBot, but any script can send that string. To confirm a request really came from Anthropic, check its source IP address against claude.com/crawling/bots.json. Anthropic says a request from an IP address on that list comes from Anthropic.
Where to see ClaudeBot visits
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so ClaudeBot won't appear there. Your server and CDN logs record every request with its user agent and IP address. Search them for ClaudeBot, Claude-User and Claude-SearchBot to see which pages each one requested and when.
OneLence AI visibility does this for you when your server reports page requests to OneLence. It matches each request to the crawler that sent it, marks the ones confirmed by Anthropic's IP list as verified, and shows which pages ClaudeBot, Claude-SearchBot and Claude-User read, next to the visitors Claude sends you and the ones that convert.
Frequently asked questions
What is ClaudeBot?
ClaudeBot is Anthropic's web crawler for training data. It collects public web content that could contribute to training Anthropic's Claude models. Anthropic runs two other agents with different jobs: Claude-User fetches pages when someone asks Claude a question, and Claude-SearchBot indexes pages for Claude's search results.
Does ClaudeBot respect robots.txt?
Yes. Anthropic says its bots honor industry standard robots.txt directives, including Crawl-delay, and don't try to get around CAPTCHAs or other anti-bot measures. A Disallow rule for ClaudeBot tells Anthropic to exclude your pages from future training data.
Will blocking ClaudeBot remove my site from Claude's answers?
Not by itself. ClaudeBot, Claude-User and Claude-SearchBot each have their own robots.txt token, so a rule for ClaudeBot doesn't apply to the other two. Blocking Claude-User stops Claude from retrieving your pages when people ask it questions, and blocking Claude-SearchBot can make your pages less visible and accurate in Claude's search results.
How do I block ClaudeBot?
Add a group to the robots.txt of every host you want covered, including subdomains: a line with User-agent: ClaudeBot, then Disallow: /. To slow it down instead of blocking it, use Crawl-delay in the same group. Anthropic advises against blocking its IP addresses, because the bots may then be unable to read your robots.txt.
How can I tell whether a ClaudeBot request is real?
Check the source IP address against the list Anthropic publishes at claude.com/crawling/bots.json. Anthropic says a request from an IP address on that list comes from Anthropic. Anyone can put ClaudeBot in a user agent, so the name alone proves nothing.
Why doesn't ClaudeBot show up in Google Analytics?
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Look for ClaudeBot in your server or CDN logs instead, where every request is recorded with its user agent and IP address.
