AI crawler guide

CCBot: Common Crawl's crawler, and what blocking it changes

CCBot builds Common Crawl's free web archive, which Common Crawl calls one of the most widely used sources of training data for large language models. It follows robots.txt and Crawl-delay, but blocking it only stops future crawls.

Last checked against official documentation on October 5, 2026.

Common Crawl is a nonprofit that crawls the web and gives its archive away. It publishes a crawl about once a month, each typically with more than two billion web pages, and says the archive has become one of the most widely used sources of training data for large language models. CCBot is the crawler that collects it.

CCBot at a glance

DetailCCBot
OperatorCommon Crawl Foundation, a 501(c)(3) nonprofit
User agentCCBot/2.0 (https://commoncrawl.org/faq/)
robots.txt nameCCBot
Follows robots.txtYes, including Crawl-delay
IP listccbot.json
Reverse DNSHosts under crawl.commoncrawl.org, except over IPv6

How CCBot treats robots.txt

Common Crawl says:

  • A CCBot disallow in your robots.txt stops its crawler crawling your site.
  • It obeys Crawl-delay: a larger number tells CCBot to slow down.
  • It honors nofollow on the links on your site.
  • It periodically checks whether your robots.txt has changed, without saying how often.
  • It slows down on its own when your server answers with HTTP 429 or 5xx errors.

robots.txt rules for CCBot

Block it:

TEXT
User-agent: CCBot
Disallow: /

Slow it down instead, as in Common Crawl's example:

TEXT
User-agent: CCBot
Crawl-delay: 2

A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.

What blocking CCBot changes, and what it doesn't

  • It stops future crawls. That's the effect Common Crawl documents.
  • It doesn't remove pages already archived. Common Crawl doesn't document removing pages from crawls it has published, and copies that others downloaded aren't affected either.
  • Legal requests are a separate route. Common Crawl says it processes legal requests as it receives them, and publishes the opt-out requests in an Opt-out Registry to alert the people who use its data.

Common Crawl argues the other side on its own blog: "If you're blocked at the crawl layer, you're excluded entirely." That's its view, and the choice is yours.

How to verify CCBot

Common Crawl runs CCBot on dedicated IP ranges with reverse DNS, except over IPv6, where reverse DNS isn't supported yet. A genuine request's IP address resolves to a host such as 18-97-14-84.crawl.commoncrawl.org. Common Crawl's own examples:

TEXT
host 18.97.14.84
dig -x 18.97.14.84
dig 18-97-14-84.crawl.commoncrawl.org A

You can also check the address against ccbot.json. Common Crawl warns that other crawlers falsely identify themselves as CCBot, so don't trust the user agent alone.

Where to see CCBot

Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so CCBot won't appear there. Your server and CDN logs record every request with its user agent and IP address.

OneLence AI visibility recognizes CCBot when your server reports page requests to OneLence, marks the requests confirmed by Common Crawl's IP list as verified, and shows which pages it read and how often, next to the crawlers of OpenAI, Anthropic, Perplexity and Google.

Frequently asked questions

What is CCBot?

CCBot is the web crawler of Common Crawl, a nonprofit that crawls the web and freely provides its archives and datasets to the public. Its user agent is CCBot/2.0 (https://commoncrawl.org/faq/).

Is Common Crawl used to train AI models?

Yes. Common Crawl says its archive has been cited in over 12,000 research papers and has become one of the most widely used sources of training data for large language models. The archive is free to download, so anyone can use it.

Does CCBot respect robots.txt?

Yes. Common Crawl says adding User-agent: CCBot and Disallow: / to your robots.txt stops its crawler, that it obeys Crawl-delay and that it honors nofollow on links. It also slows down when your server answers with HTTP 429 or 5xx errors.

Does blocking CCBot remove my pages from Common Crawl?

Blocking stops future crawls. Common Crawl doesn't document removing pages from crawls it has already published, and copies that others have downloaded aren't affected. It processes legal requests and publishes the opt-out requests it receives in an Opt-out Registry.

How do I slow CCBot down?

Add Crawl-delay to a CCBot group in your robots.txt, as in Common Crawl's own example: User-agent: CCBot, then Crawl-delay: 2. A larger number tells CCBot to crawl more slowly.

How can I tell whether a request really came from Common Crawl?

Run a reverse DNS lookup on the source IP address: CCBot runs on dedicated IP ranges whose hosts are under crawl.commoncrawl.org, except over IPv6. Or check the address against index.commoncrawl.org/ccbot.json. Common Crawl warns that other crawlers falsely identify themselves as CCBot.