Common Crawl is a nonprofit that crawls the web and gives its archive away. It publishes a crawl about once a month, each typically with more than two billion web pages, and says the archive has become one of the most widely used sources of training data for large language models. CCBot is the crawler that collects it.
CCBot at a glance
| Detail | CCBot |
|---|---|
| Operator | Common Crawl Foundation, a 501(c)(3) nonprofit |
| User agent | CCBot/2.0 (https://commoncrawl.org/faq/) |
| robots.txt name | CCBot |
| Follows robots.txt | Yes, including Crawl-delay |
| IP list | ccbot.json |
| Reverse DNS | Hosts under crawl.commoncrawl.org, except over IPv6 |
How CCBot treats robots.txt
Common Crawl says:
- A CCBot disallow in your robots.txt stops its crawler crawling your site.
- It obeys
Crawl-delay: a larger number tells CCBot to slow down. - It honors
nofollowon the links on your site. - It periodically checks whether your robots.txt has changed, without saying how often.
- It slows down on its own when your server answers with HTTP 429 or 5xx errors.
robots.txt rules for CCBot
Block it:
User-agent: CCBot
Disallow: /
Slow it down instead, as in Common Crawl's example:
User-agent: CCBot
Crawl-delay: 2
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
What blocking CCBot changes, and what it doesn't
- It stops future crawls. That's the effect Common Crawl documents.
- It doesn't remove pages already archived. Common Crawl doesn't document removing pages from crawls it has published, and copies that others downloaded aren't affected either.
- Legal requests are a separate route. Common Crawl says it processes legal requests as it receives them, and publishes the opt-out requests in an Opt-out Registry to alert the people who use its data.
Common Crawl argues the other side on its own blog: "If you're blocked at the crawl layer, you're excluded entirely." That's its view, and the choice is yours.
How to verify CCBot
Common Crawl runs CCBot on dedicated IP ranges with reverse DNS, except over IPv6, where reverse DNS isn't supported yet. A genuine request's IP address resolves to a host such as 18-97-14-84.crawl.commoncrawl.org. Common Crawl's own examples:
host 18.97.14.84
dig -x 18.97.14.84
dig 18-97-14-84.crawl.commoncrawl.org A
You can also check the address against ccbot.json. Common Crawl warns that other crawlers falsely identify themselves as CCBot, so don't trust the user agent alone.
Where to see CCBot
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so CCBot won't appear there. Your server and CDN logs record every request with its user agent and IP address.
OneLence AI visibility recognizes CCBot when your server reports page requests to OneLence, marks the requests confirmed by Common Crawl's IP list as verified, and shows which pages it read and how often, next to the crawlers of OpenAI, Anthropic, Perplexity and Google.
Frequently asked questions
What is CCBot?
CCBot is the web crawler of Common Crawl, a nonprofit that crawls the web and freely provides its archives and datasets to the public. Its user agent is CCBot/2.0 (https://commoncrawl.org/faq/).
Is Common Crawl used to train AI models?
Yes. Common Crawl says its archive has been cited in over 12,000 research papers and has become one of the most widely used sources of training data for large language models. The archive is free to download, so anyone can use it.
Does CCBot respect robots.txt?
Yes. Common Crawl says adding User-agent: CCBot and Disallow: / to your robots.txt stops its crawler, that it obeys Crawl-delay and that it honors nofollow on links. It also slows down when your server answers with HTTP 429 or 5xx errors.
Does blocking CCBot remove my pages from Common Crawl?
Blocking stops future crawls. Common Crawl doesn't document removing pages from crawls it has already published, and copies that others have downloaded aren't affected. It processes legal requests and publishes the opt-out requests it receives in an Opt-out Registry.
How do I slow CCBot down?
Add Crawl-delay to a CCBot group in your robots.txt, as in Common Crawl's own example: User-agent: CCBot, then Crawl-delay: 2. A larger number tells CCBot to crawl more slowly.
How can I tell whether a request really came from Common Crawl?
Run a reverse DNS lookup on the source IP address: CCBot runs on dedicated IP ranges whose hosts are under crawl.commoncrawl.org, except over IPv6. Or check the address against index.commoncrawl.org/ccbot.json. Common Crawl warns that other crawlers falsely identify themselves as CCBot.
