[{"data":1,"prerenderedAt":3358},["ShallowReactive",2],{"ai-crawlers:\u002Fai-crawlers\u002Fclaudebot":3},{"page":4,"pages":235},{"id":5,"title":6,"body":7,"checked":202,"description":203,"documented":204,"extension":205,"faq":206,"meta":225,"name":49,"navigation":204,"order":226,"path":227,"seo":228,"sitemap":231,"stem":232,"tagline":233,"__hash__":234},"aiCrawlers\u002Fai-crawlers\u002Fclaudebot.md","ClaudeBot: what Anthropic's crawler does and how to control it",{"type":8,"value":9,"toc":194},"minimark",[10,14,19,82,93,97,104,107,111,114,124,127,133,136,142,157,161,171,175,187],[11,12,13],"p",{},"ClaudeBot is one of three web agents Anthropic runs. Each has its own job and its own robots.txt token, so you can allow one and block another.",[15,16,18],"h2",{"id":17},"anthropics-three-bots","Anthropic's three bots",[20,21,22,38],"table",{},[23,24,25],"thead",{},[26,27,28,32,35],"tr",{},[29,30,31],"th",{},"Bot",[29,33,34],{},"What it does",[29,36,37],{},"What blocking it in robots.txt changes",[39,40,41,56,69],"tbody",{},[26,42,43,50,53],{},[44,45,46],"td",{},[47,48,49],"code",{},"ClaudeBot",[44,51,52],{},"Collects public web content that could contribute to training Claude models",[44,54,55],{},"Tells Anthropic to exclude your pages from future training data",[26,57,58,63,66],{},[44,59,60],{},[47,61,62],{},"Claude-User",[44,64,65],{},"Visits a page when someone asks Claude a question that needs it",[44,67,68],{},"Claude can't retrieve your pages to answer people's questions",[26,70,71,76,79],{},[44,72,73],{},[47,74,75],{},"Claude-SearchBot",[44,77,78],{},"Crawls the web to improve the quality of Claude's search results",[44,80,81],{},"Your pages can be less visible and less accurate in Claude's search results",[11,83,84,85,92],{},"Anthropic publishes the IP addresses its bots use in one list, ",[86,87,91],"a",{"href":88,"rel":89},"https:\u002F\u002Fclaude.com\u002Fcrawling\u002Fbots.json",[90],"nofollow","claude.com\u002Fcrawling\u002Fbots.json",".",[15,94,96],{"id":95},"does-claudebot-follow-robotstxt","Does ClaudeBot follow robots.txt?",[11,98,99,100,103],{},"Yes. Anthropic says its bots honor industry standard robots.txt directives, support ",[47,101,102],{},"Crawl-delay",", and don't try to get around CAPTCHAs or other anti-bot measures. robots.txt works per host, so add the rules to the robots.txt of each subdomain you want covered.",[11,105,106],{},"Anthropic advises against blocking its bots by IP address. It may not work as intended, and it can stop the bots from reading your robots.txt at all.",[15,108,110],{"id":109},"robotstxt-rules-for-claudebot","robots.txt rules for ClaudeBot",[11,112,113],{},"Block ClaudeBot from the whole site:",[115,116,121],"pre",{"className":117,"code":119,"language":120},[118],"language-text","User-agent: ClaudeBot\nDisallow: \u002F\n","text",[47,122,119],{"__ignoreMap":123},"",[11,125,126],{},"Slow it down instead of blocking it:",[115,128,131],{"className":129,"code":130,"language":120},[118],"User-agent: ClaudeBot\nCrawl-delay: 1\n",[47,132,130],{"__ignoreMap":123},[11,134,135],{},"Opt out of training while keeping Claude's answers and search:",[115,137,140],{"className":138,"code":139,"language":120},[118],"User-agent: ClaudeBot\nDisallow: \u002F\n\nUser-agent: Claude-User\nAllow: \u002F\n\nUser-agent: Claude-SearchBot\nAllow: \u002F\n",[47,141,139],{"__ignoreMap":123},[11,143,144,145,148,149,152,153,156],{},"A crawler that finds a group with its own name follows only that group and ignores the ",[47,146,147],{},"User-agent: *"," group. If your ",[47,150,151],{},"*"," group disallows paths such as ",[47,154,155],{},"\u002Fadmin\u002F",", repeat those lines in each named group that should still respect them.",[15,158,160],{"id":159},"how-to-verify-a-claudebot-request","How to verify a ClaudeBot request",[11,162,163,164,166,167,170],{},"The user agent of a ClaudeBot request contains ",[47,165,49],{},", but any script can send that string. To confirm a request really came from Anthropic, check its source IP address against ",[86,168,91],{"href":88,"rel":169},[90],". Anthropic says a request from an IP address on that list comes from Anthropic.",[15,172,174],{"id":173},"where-to-see-claudebot-visits","Where to see ClaudeBot visits",[11,176,177,178,180,181,183,184,186],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so ClaudeBot won't appear there. Your server and CDN logs record every request with its user agent and IP address. Search them for ",[47,179,49],{},", ",[47,182,62],{}," and ",[47,185,75],{}," to see which pages each one requested and when.",[11,188,189,193],{},[86,190,192],{"href":191},"\u002Fai-visibility","OneLence AI visibility"," does this for you when your server reports page requests to OneLence. It matches each request to the crawler that sent it, marks the ones confirmed by Anthropic's IP list as verified, and shows which pages ClaudeBot, Claude-SearchBot and Claude-User read, next to the visitors Claude sends you and the ones that convert.",{"title":123,"searchDepth":195,"depth":195,"links":196},2,[197,198,199,200,201],{"id":17,"depth":195,"text":18},{"id":95,"depth":195,"text":96},{"id":109,"depth":195,"text":110},{"id":159,"depth":195,"text":160},{"id":173,"depth":195,"text":174},"2026-10-04","ClaudeBot is the crawler Anthropic uses to collect public web content that may be used to train Claude. What it does, how it differs from Claude-User and Claude-SearchBot, and how to block, slow down or verify it.",true,"md",[207,210,213,216,219,222],{"question":208,"answer":209},"What is ClaudeBot?","ClaudeBot is Anthropic's web crawler for training data. It collects public web content that could contribute to training Anthropic's Claude models. Anthropic runs two other agents with different jobs: Claude-User fetches pages when someone asks Claude a question, and Claude-SearchBot indexes pages for Claude's search results.",{"question":211,"answer":212},"Does ClaudeBot respect robots.txt?","Yes. Anthropic says its bots honor industry standard robots.txt directives, including Crawl-delay, and don't try to get around CAPTCHAs or other anti-bot measures. A Disallow rule for ClaudeBot tells Anthropic to exclude your pages from future training data.",{"question":214,"answer":215},"Will blocking ClaudeBot remove my site from Claude's answers?","Not by itself. ClaudeBot, Claude-User and Claude-SearchBot each have their own robots.txt token, so a rule for ClaudeBot doesn't apply to the other two. Blocking Claude-User stops Claude from retrieving your pages when people ask it questions, and blocking Claude-SearchBot can make your pages less visible and accurate in Claude's search results.",{"question":217,"answer":218},"How do I block ClaudeBot?","Add a group to the robots.txt of every host you want covered, including subdomains: a line with User-agent: ClaudeBot, then Disallow: \u002F. To slow it down instead of blocking it, use Crawl-delay in the same group. Anthropic advises against blocking its IP addresses, because the bots may then be unable to read your robots.txt.",{"question":220,"answer":221},"How can I tell whether a ClaudeBot request is real?","Check the source IP address against the list Anthropic publishes at claude.com\u002Fcrawling\u002Fbots.json. Anthropic says a request from an IP address on that list comes from Anthropic. Anyone can put ClaudeBot in a user agent, so the name alone proves nothing.",{"question":223,"answer":224},"Why doesn't ClaudeBot show up in Google Analytics?","Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Look for ClaudeBot in your server or CDN logs instead, where every request is recorded with its user agent and IP address.",{},1,"\u002Fai-crawlers\u002Fclaudebot",{"title":229,"description":230},"ClaudeBot: Anthropic's Crawler Explained, and How to Block It","What ClaudeBot does, how it differs from Claude-User and Claude-SearchBot, how to check its IP addresses, and the robots.txt rules that block, slow down or allow it.",{"loc":227},"ai-crawlers\u002Fclaudebot","ClaudeBot collects public web content that may be used to train Anthropic's Claude models. It follows robots.txt, and a rule for ClaudeBot doesn't apply to Claude-User or Claude-SearchBot, the agents behind Claude's answers and search.","tacNU8XSTIyd-WQvRncPvi3bDZIMp6-x8xAPqHT5f-I",[236,1064,1199,1428,1771,1979,2252,2497,2731,2950,3116],{"id":237,"title":238,"body":239,"checked":202,"description":1033,"documented":204,"extension":205,"faq":1034,"meta":1053,"name":1054,"navigation":204,"order":1055,"path":1056,"seo":1057,"sitemap":1060,"stem":1061,"tagline":1062,"__hash__":1063},"aiCrawlers\u002Fai-crawlers\u002Findex.md","AI crawlers: who they are, what they do and how to control them",{"type":8,"value":240,"toc":1024},[241,245,248,298,306,310,815,818,821,887,904,908,941,945,948,954,957,977,989,993,996,999,1003,1009,1013,1016,1019],[15,242,244],{"id":243},"three-jobs-three-kinds-of-crawler","Three jobs, three kinds of crawler",[11,246,247],{},"AI companies don't send one crawler. They send different agents for different jobs, and blocking one doesn't block the others.",[20,249,250,263],{},[23,251,252],{},[26,253,254,257,260],{},[29,255,256],{},"Job",[29,258,259],{},"What the crawler does",[29,261,262],{},"Examples",[39,264,265,276,287],{},[26,266,267,270,273],{},[44,268,269],{},"Training",[44,271,272],{},"Collects pages that may be used to train future models",[44,274,275],{},"GPTBot, ClaudeBot, Meta-ExternalAgent, MistralAI-Training",[26,277,278,281,284],{},[44,279,280],{},"AI search",[44,282,283],{},"Builds the index that AI search and answers draw on",[44,285,286],{},"OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer",[26,288,289,292,295],{},[44,290,291],{},"Answers",[44,293,294],{},"Fetches a page while an assistant answers someone, because their request needs it",[44,296,297],{},"ChatGPT-User, Claude-User, Perplexity-User, Google-Agent, DuckAssistBot",[11,299,300,301,305],{},"Google works differently. AI Overviews and AI Mode are part of Google Search, so robots.txt rules for Googlebot control how they crawl, and since August 31, 2026, a Search Console setting can opt a site out of them. Google-Extended isn't a crawler at all, but a robots.txt name that controls whether pages Google already crawls may be used to train Gemini models and to ground Gemini's answers. The ",[86,302,304],{"href":303},"\u002Fai-crawlers\u002Fgoogle-extended","Google-Extended guide"," covers both.",[15,307,309],{"id":308},"the-main-ai-crawlers","The main AI crawlers",[20,311,312,330],{},[23,313,314],{},[26,315,316,319,322,324,327],{},[29,317,318],{},"Company",[29,320,321],{},"Agent",[29,323,256],{},[29,325,326],{},"Follows robots.txt",[29,328,329],{},"IP list",[39,331,332,354,375,396,415,433,450,472,493,511,533,556,577,599,621,638,656,672,689,708,731,754,777,799],{},[26,333,334,337,342,344,347],{},[44,335,336],{},"OpenAI",[44,338,339],{},[47,340,341],{},"GPTBot",[44,343,269],{},[44,345,346],{},"Yes",[44,348,349],{},[86,350,353],{"href":351,"rel":352},"https:\u002F\u002Fopenai.com\u002Fgptbot.json",[90],"gptbot.json",[26,355,356,358,363,366,368],{},[44,357,336],{},[44,359,360],{},[47,361,362],{},"OAI-SearchBot",[44,364,365],{},"ChatGPT search",[44,367,346],{},[44,369,370],{},[86,371,374],{"href":372,"rel":373},"https:\u002F\u002Fopenai.com\u002Fsearchbot.json",[90],"searchbot.json",[26,376,377,379,384,386,389],{},[44,378,336],{},[44,380,381],{},[47,382,383],{},"ChatGPT-User",[44,385,291],{},[44,387,388],{},"May not apply, says OpenAI",[44,390,391],{},[86,392,395],{"href":393,"rel":394},"https:\u002F\u002Fopenai.com\u002Fchatgpt-user.json",[90],"chatgpt-user.json",[26,397,398,401,405,407,409],{},[44,399,400],{},"Anthropic",[44,402,403],{},[47,404,49],{},[44,406,269],{},[44,408,346],{},[44,410,411],{},[86,412,414],{"href":88,"rel":413},[90],"bots.json",[26,416,417,419,423,426,428],{},[44,418,400],{},[44,420,421],{},[47,422,75],{},[44,424,425],{},"Claude search",[44,427,346],{},[44,429,430],{},[86,431,414],{"href":88,"rel":432},[90],[26,434,435,437,441,443,445],{},[44,436,400],{},[44,438,439],{},[47,440,62],{},[44,442,291],{},[44,444,346],{},[44,446,447],{},[86,448,414],{"href":88,"rel":449},[90],[26,451,452,455,460,463,465],{},[44,453,454],{},"Perplexity",[44,456,457],{},[47,458,459],{},"PerplexityBot",[44,461,462],{},"Perplexity search, not model training",[44,464,346],{},[44,466,467],{},[86,468,471],{"href":469,"rel":470},"https:\u002F\u002Fwww.perplexity.com\u002Fperplexitybot.json",[90],"perplexitybot.json",[26,473,474,476,481,483,486],{},[44,475,454],{},[44,477,478],{},[47,479,480],{},"Perplexity-User",[44,482,291],{},[44,484,485],{},"Generally not, says Perplexity",[44,487,488],{},[86,489,492],{"href":490,"rel":491},"https:\u002F\u002Fwww.perplexity.com\u002Fperplexity-user.json",[90],"perplexity-user.json",[26,494,495,498,503,506,508],{},[44,496,497],{},"Google",[44,499,500],{},[47,501,502],{},"Google-Extended",[44,504,505],{},"Gemini training and grounding (a robots.txt name, not a crawler)",[44,507,346],{},[44,509,510],{},"Uses Google's existing crawlers",[26,512,513,515,520,523,526],{},[44,514,497],{},[44,516,517],{},[47,518,519],{},"Google-Agent",[44,521,522],{},"Agents acting on a person's request",[44,524,525],{},"Generally not, says Google",[44,527,528],{},[86,529,532],{"href":530,"rel":531},"https:\u002F\u002Fdevelopers.google.com\u002Fstatic\u002Fcrawling\u002Fipranges\u002Fuser-triggered-agents.json",[90],"user-triggered-agents.json",[26,534,535,538,543,546,549],{},[44,536,537],{},"Amazon",[44,539,540],{},[47,541,542],{},"Amazonbot",[44,544,545],{},"Improving Amazon's products; may train Amazon AI models",[44,547,548],{},"Yes, but not Crawl-delay",[44,550,551],{},[86,552,555],{"href":553,"rel":554},"https:\u002F\u002Fdeveloper.amazon.com\u002Famazonbot\u002Fip-addresses\u002F",[90],"Amazonbot list",[26,557,558,560,565,568,570],{},[44,559,537],{},[44,561,562],{},[47,563,564],{},"Amzn-SearchBot",[44,566,567],{},"Search experiences such as Alexa, not training",[44,569,548],{},[44,571,572],{},[86,573,576],{"href":574,"rel":575},"https:\u002F\u002Fdeveloper.amazon.com\u002Famazonbot\u002Fsearchbot-ip-addresses\u002F",[90],"Amzn-SearchBot list",[26,578,579,581,586,589,592],{},[44,580,537],{},[44,582,583],{},[47,584,585],{},"Amzn-User",[44,587,588],{},"Answers, such as Alexa questions",[44,590,591],{},"May not follow every directive, says Amazon",[44,593,594],{},[86,595,598],{"href":596,"rel":597},"https:\u002F\u002Fdeveloper.amazon.com\u002Famazonbot\u002Flive-ip-addresses\u002F",[90],"Amzn-User list",[26,600,601,604,609,612,614],{},[44,602,603],{},"Apple",[44,605,606],{},[47,607,608],{},"Applebot",[44,610,611],{},"Search in Siri, Spotlight and Safari; may train Apple models",[44,613,548],{},[44,615,616],{},[86,617,620],{"href":618,"rel":619},"https:\u002F\u002Fsearch.developer.apple.com\u002Fapplebot.json",[90],"applebot.json",[26,622,623,625,630,633,635],{},[44,624,603],{},[44,626,627],{},[47,628,629],{},"Applebot-Extended",[44,631,632],{},"Apple model training (a robots.txt name, not a crawler)",[44,634,346],{},[44,636,637],{},"Uses Applebot's addresses",[26,639,640,643,648,651,653],{},[44,641,642],{},"Meta",[44,644,645],{},[47,646,647],{},"Meta-ExternalAgent",[44,649,650],{},"Uses such as AI model training and direct indexing",[44,652,346],{},[44,654,655],{},"Meta's AS32934, via whois",[26,657,658,660,665,668,670],{},[44,659,642],{},[44,661,662],{},[47,663,664],{},"Meta-WebIndexer",[44,666,667],{},"Meta AI search and citations",[44,669,346],{},[44,671,655],{},[26,673,674,676,681,684,687],{},[44,675,642],{},[44,677,678],{},[47,679,680],{},"Meta-ExternalFetcher",[44,682,683],{},"Fetches for users and AI agent tasks",[44,685,686],{},"May bypass it, says Meta",[44,688,655],{},[26,690,691,694,699,702,705],{},[44,692,693],{},"ByteDance",[44,695,696],{},[47,697,698],{},"Bytespider",[44,700,701],{},"Not documented by ByteDance",[44,703,704],{},"Not reliably, say independent reports",[44,706,707],{},"None published",[26,709,710,713,718,721,724],{},[44,711,712],{},"Common Crawl",[44,714,715],{},[47,716,717],{},"CCBot",[44,719,720],{},"Open web archive, widely used to train AI models",[44,722,723],{},"Yes, including Crawl-delay",[44,725,726],{},[86,727,730],{"href":728,"rel":729},"https:\u002F\u002Findex.commoncrawl.org\u002Fccbot.json",[90],"ccbot.json",[26,732,733,736,741,744,747],{},[44,734,735],{},"DuckDuckGo",[44,737,738],{},[47,739,740],{},"DuckAssistBot",[44,742,743],{},"Answers in DuckAssist, not training",[44,745,746],{},"Yes, within 72 hours",[44,748,749],{},[86,750,753],{"href":751,"rel":752},"https:\u002F\u002Fduckduckgo.com\u002Fduckduckgo-help-pages\u002Fresults\u002Fduckassistbot",[90],"On its help page",[26,755,756,759,764,767,770],{},[44,757,758],{},"Mistral",[44,760,761],{},[47,762,763],{},"MistralAI-User",[44,765,766],{},"Answers in Vibe, formerly Le Chat",[44,768,769],{},"Mistral says it governs which sites it fetches",[44,771,772],{},[86,773,776],{"href":774,"rel":775},"https:\u002F\u002Fmistral.ai\u002Fmistralai-user-ips.json",[90],"mistralai-user-ips.json",[26,778,779,781,786,789,792],{},[44,780,758],{},[44,782,783],{},[47,784,785],{},"MistralAI-Index",[44,787,788],{},"Mistral search, not training",[44,790,791],{},"Not stated",[44,793,794],{},[86,795,798],{"href":796,"rel":797},"https:\u002F\u002Fmistral.ai\u002Fmistralai-index-ips.json",[90],"mistralai-index-ips.json",[26,800,801,803,808,811,813],{},[44,802,758],{},[44,804,805],{},[47,806,807],{},"MistralAI-Training",[44,809,810],{},"Training Mistral models",[44,812,346],{},[44,814,707],{},[11,816,817],{},"Perplexity adds one detail: when robots.txt blocks PerplexityBot, Perplexity may still index the page's domain, headline and a brief factual summary.",[11,819,820],{},"Guides to the most searched-for ones:",[822,823,824,830,837,844,849,855,862,868,874,880],"ul",{},[825,826,827,829],"li",{},[86,828,49],{"href":227},": Anthropic's training crawler, and how it differs from Claude-User and Claude-SearchBot.",[825,831,832,836],{},[86,833,835],{"href":834},"\u002Fai-crawlers\u002Fchatgpt-user","ChatGPT-User vs GPTBot vs OAI-SearchBot",": OpenAI's three agents, and which one decides whether you appear in ChatGPT search.",[825,838,839,843],{},[86,840,842],{"href":841},"\u002Fai-crawlers\u002Fperplexitybot","PerplexityBot and Perplexity-User",": Perplexity's search crawler and the agent that fetches pages for answers.",[825,845,846,848],{},[86,847,502],{"href":303},": what it controls, and the Search Console setting for AI Overviews and AI Mode.",[825,850,851,854],{},[86,852,542],{"href":853},"\u002Fai-crawlers\u002Famazonbot",": Amazon's three agents, and how to opt out of training but stay in Alexa.",[825,856,857,861],{},[86,858,860],{"href":859},"\u002Fai-crawlers\u002Fapplebot","Applebot and Applebot-Extended",": Apple's crawler and its training opt-out.",[825,863,864,867],{},[86,865,647],{"href":866},"\u002Fai-crawlers\u002Fmeta-externalagent",": Meta's five crawlers, and how to block training without breaking link previews.",[825,869,870,873],{},[86,871,698],{"href":872},"\u002Fai-crawlers\u002Fbytespider",": ByteDance's crawler, and why robots.txt may not stop it.",[825,875,876,879],{},[86,877,717],{"href":878},"\u002Fai-crawlers\u002Fccbot",": Common Crawl's crawler, and what blocking it changes.",[825,881,882,886],{},[86,883,885],{"href":884},"\u002Fai-crawlers\u002Fllms-txt","llms.txt",": the proposed file for AI agents, and who actually reads it.",[11,888,889,890,894,895,899,900,92],{},"To see which of these crawlers your robots.txt lets read a page, and the line that decides it, use the free ",[86,891,893],{"href":892},"\u002Fai-crawlers\u002Frobots-txt-checker","robots.txt checker for AI crawlers",". To write a robots.txt that allows or blocks each of them, use the free ",[86,896,898],{"href":897},"\u002Fai-crawlers\u002Frobots-txt-generator","robots.txt generator",", and to draft an llms.txt from your own pages, the free ",[86,901,903],{"href":902},"\u002Fai-crawlers\u002Fllms-txt-generator","llms.txt generator",[15,905,907],{"id":906},"other-ai-agents-you-may-see","Other AI agents you may see",[822,909,910,917,923,929,935],{},[825,911,912,916],{},[913,914,915],"strong",{},"Diffbot"," crawls for Diffbot's Knowledge Graph and web search, which Diffbot says isn't used for AI training. It follows robots.txt, including Crawl-delay, by default, but Diffbot's customers can switch that off for crawls they run with its software and can send their own user agent.",[825,918,919,922],{},[913,920,921],{},"KimiBot, Kimi-SearchBot and Kimi-User"," (Moonshot AI) collect training data, build Kimi's search and fetch pages for users. Moonshot publishes an IP list for each, and says robots.txt rules may not directly apply to Kimi-User.",[825,924,925,928],{},[913,926,927],{},"AI2Bot"," (Allen Institute for AI) collects web content used to train open language models.",[825,930,931,934],{},[913,932,933],{},"YouBot"," (You.com) is no longer documented by You.com: about.you.com\u002Fyoubot, the page crawler directories cite for it, returns a 404 error, so requests that use the name can't be verified.",[825,936,937,940],{},[913,938,939],{},"xAI's Grok and DeepSeek"," publish no crawler documentation we could find.",[15,942,944],{"id":943},"opt-out-of-training-without-leaving-ai-answers","Opt out of training without leaving AI answers",[11,946,947],{},"These rules ask OpenAI, Anthropic and Google not to use your pages for model training, and leave ChatGPT search, Claude's search and the answers that fetch your pages alone:",[115,949,952],{"className":950,"code":951,"language":120},[118],"User-agent: GPTBot\nDisallow: \u002F\n\nUser-agent: ClaudeBot\nDisallow: \u002F\n\nUser-agent: Google-Extended\nDisallow: \u002F\n\nUser-agent: Applebot-Extended\nDisallow: \u002F\n\nUser-agent: MistralAI-Training\nDisallow: \u002F\n",[47,953,951],{"__ignoreMap":123},[11,955,956],{},"Two things to know before you copy them:",[822,958,959,965,971],{},[825,960,961,964],{},[913,962,963],{},"Google-Extended does more than training."," Google says it also controls grounding, the content Gemini Apps and Grounding with Google Search on Vertex AI pull from Google's index while answering. Blocking it doesn't affect Google Search, AI Overviews or AI Mode.",[825,966,967,970],{},[913,968,969],{},"Perplexity has no training crawler to block."," Perplexity says it doesn't build foundation models, so your content isn't used for pre-training. PerplexityBot only feeds Perplexity's search.",[825,972,973,976],{},[913,974,975],{},"Some crawlers do more than training."," Amazonbot, Meta-ExternalAgent and CCBot collect data that may be used for training, but each serves other purposes too, so read their guides before you block them.",[11,978,144,979,148,981,152,983,985,986,988],{},[47,980,147],{},[47,982,151],{},[47,984,155],{},", repeat those lines in each named group that should still respect them. The free ",[86,987,898],{"href":897}," writes these groups for you and repeats your private paths in each group that lets a crawler in.",[15,990,992],{"id":991},"agents-that-act-for-a-person","Agents that act for a person",[11,994,995],{},"The agents that fetch a page because someone asked for it often don't follow robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says its user-triggered fetchers, including Google-Agent, generally ignore them, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic is the exception: it says its bots honor robots.txt, and that blocking Claude-User stops Claude from retrieving your pages for people's questions.",[11,997,998],{},"Keeping these agents out takes a firewall rule rather than robots.txt, and it also keeps your pages out of the answers people ask for. Think about that trade before you block them.",[15,1000,1002],{"id":1001},"how-to-verify-an-ai-crawler","How to verify an AI crawler",[11,1004,1005,1006,92],{},"User agent strings are easy to fake, so a request that says GPTBot or ClaudeBot proves nothing on its own. Every company in the table above publishes the IP addresses its agents use. Check the request's source IP against the right list: when it matches, the request is real. For Google's common crawlers, Google also documents a reverse DNS check: their hostnames match ",[47,1007,1008],{},"crawl-***-***-***-***.googlebot.com",[15,1010,1012],{"id":1011},"where-to-see-ai-crawlers","Where to see AI crawlers",[11,1014,1015],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up. Search them for the agent names in the table above.",[11,1017,1018],{},"Make sure the content you want read is in the HTML your server sends. Google says it can process content in JavaScript as long as it isn't blocked, but OpenAI, Anthropic and Perplexity don't say whether their crawlers run JavaScript, and a page that only fills in after scripts run may look empty to them.",[11,1020,1021,1023],{},[86,1022,192],{"href":191}," turns those requests into a report when your server reports them to OneLence. It matches each request to the crawler that sent it, marks the ones confirmed by the company's IP list as verified, and shows which pages each crawler read, which pages assistants fetched while answering people, and the visitors those assistants sent you and their conversions.",{"title":123,"searchDepth":195,"depth":195,"links":1025},[1026,1027,1028,1029,1030,1031,1032],{"id":243,"depth":195,"text":244},{"id":308,"depth":195,"text":309},{"id":906,"depth":195,"text":907},{"id":943,"depth":195,"text":944},{"id":991,"depth":195,"text":992},{"id":1001,"depth":195,"text":1002},{"id":1011,"depth":195,"text":1012},"A guide to the crawlers OpenAI, Anthropic, Perplexity, Google, Amazon, Apple, Meta and others send to websites: which ones train models, which ones feed AI search and answers, which follow robots.txt, and how to verify them.",[1035,1038,1041,1044,1047,1050],{"question":1036,"answer":1037},"What is an AI crawler?","An AI crawler is an automated client that an AI company sends to websites. Some collect pages to train models (GPTBot, ClaudeBot), some build the index behind AI search (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and some fetch a page while an assistant answers someone (ChatGPT-User, Claude-User, Perplexity-User).",{"question":1039,"answer":1040},"Will blocking AI training crawlers remove my site from ChatGPT or Claude?","No. GPTBot and ClaudeBot only collect training data. ChatGPT search relies on OAI-SearchBot and Claude's search on Claude-SearchBot, and each has its own robots.txt name, so a rule for a training crawler doesn't apply to them. Google-Extended is the exception to keep in mind: besides Gemini training, it also controls grounding in Gemini Apps and Vertex AI.",{"question":1042,"answer":1043},"Does blocking Google-Extended remove my site from AI Overviews?","No. Google says Google-Extended doesn't affect inclusion or ranking in Google Search, and AI Overviews and AI Mode are part of Search. To opt out of them, use the Search generative AI control in Search Console, which Google rolled out to all websites worldwide as of August 31, 2026. It removes your links and content from AI Overviews, AI Mode and generative AI features in Discover without affecting the rest of Search.",{"question":1045,"answer":1046},"Which AI crawlers ignore robots.txt?","Mostly the ones that act for a person. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, Google says the same of its user-triggered fetchers such as Google-Agent, Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic says its bots, Claude-User included, honor robots.txt. Bytespider is the exception among the automatic crawlers: several independent reports say it doesn't reliably follow robots.txt.",{"question":1048,"answer":1049},"How do I know a request really comes from an AI company?","Check the source IP address against the list the company publishes, such as openai.com\u002Fgptbot.json, claude.com\u002Fcrawling\u002Fbots.json or perplexity.com\u002Fperplexitybot.json. Anyone can copy a crawler's user agent string, so the name alone proves nothing.",{"question":1051,"answer":1052},"Why don't AI crawlers show up in Google Analytics?","Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off. Server and CDN logs record every request with its user agent and IP address, so that's where AI crawlers show up.",{},"AI crawlers",0,"\u002Fai-crawlers",{"title":1058,"description":1059},"AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot and More","Which AI crawlers OpenAI, Anthropic, Perplexity and Google run, what each one does, which follow robots.txt, how to verify them and how to block training without leaving AI answers.",{"loc":1056},"ai-crawlers\u002Findex","AI companies send separate crawlers to train models, to build AI search indexes and to fetch pages while answering someone. Each has its own robots.txt name, so you can choose which ones read your site.","6efIs2OmwJOzAsXxXidswEwRQ5KkYaZxEGOuWe66pHw",{"id":5,"title":6,"body":1065,"checked":202,"description":203,"documented":204,"extension":205,"faq":1189,"meta":1196,"name":49,"navigation":204,"order":226,"path":227,"seo":1197,"sitemap":1198,"stem":232,"tagline":233,"__hash__":234},{"type":8,"value":1066,"toc":1182},[1067,1069,1071,1115,1120,1122,1126,1128,1130,1132,1137,1139,1144,1146,1151,1159,1161,1168,1170,1178],[11,1068,13],{},[15,1070,18],{"id":17},[20,1072,1073,1083],{},[23,1074,1075],{},[26,1076,1077,1079,1081],{},[29,1078,31],{},[29,1080,34],{},[29,1082,37],{},[39,1084,1085,1095,1105],{},[26,1086,1087,1091,1093],{},[44,1088,1089],{},[47,1090,49],{},[44,1092,52],{},[44,1094,55],{},[26,1096,1097,1101,1103],{},[44,1098,1099],{},[47,1100,62],{},[44,1102,65],{},[44,1104,68],{},[26,1106,1107,1111,1113],{},[44,1108,1109],{},[47,1110,75],{},[44,1112,78],{},[44,1114,81],{},[11,1116,84,1117,92],{},[86,1118,91],{"href":88,"rel":1119},[90],[15,1121,96],{"id":95},[11,1123,99,1124,103],{},[47,1125,102],{},[11,1127,106],{},[15,1129,110],{"id":109},[11,1131,113],{},[115,1133,1135],{"className":1134,"code":119,"language":120},[118],[47,1136,119],{"__ignoreMap":123},[11,1138,126],{},[115,1140,1142],{"className":1141,"code":130,"language":120},[118],[47,1143,130],{"__ignoreMap":123},[11,1145,135],{},[115,1147,1149],{"className":1148,"code":139,"language":120},[118],[47,1150,139],{"__ignoreMap":123},[11,1152,144,1153,148,1155,152,1157,156],{},[47,1154,147],{},[47,1156,151],{},[47,1158,155],{},[15,1160,160],{"id":159},[11,1162,163,1163,166,1165,170],{},[47,1164,49],{},[86,1166,91],{"href":88,"rel":1167},[90],[15,1169,174],{"id":173},[11,1171,177,1172,180,1174,183,1176,186],{},[47,1173,49],{},[47,1175,62],{},[47,1177,75],{},[11,1179,1180,193],{},[86,1181,192],{"href":191},{"title":123,"searchDepth":195,"depth":195,"links":1183},[1184,1185,1186,1187,1188],{"id":17,"depth":195,"text":18},{"id":95,"depth":195,"text":96},{"id":109,"depth":195,"text":110},{"id":159,"depth":195,"text":160},{"id":173,"depth":195,"text":174},[1190,1191,1192,1193,1194,1195],{"question":208,"answer":209},{"question":211,"answer":212},{"question":214,"answer":215},{"question":217,"answer":218},{"question":220,"answer":221},{"question":223,"answer":224},{},{"title":229,"description":230},{"loc":227},{"id":1200,"title":1201,"body":1202,"checked":202,"description":1396,"documented":204,"extension":205,"faq":1397,"meta":1419,"name":1420,"navigation":204,"order":195,"path":834,"seo":1421,"sitemap":1424,"stem":1425,"tagline":1426,"__hash__":1427},"aiCrawlers\u002Fai-crawlers\u002Fchatgpt-user.md","ChatGPT-User vs GPTBot vs OAI-SearchBot: OpenAI's crawlers explained",{"type":8,"value":1203,"toc":1387},[1204,1207,1211,1279,1291,1295,1298,1301,1305,1308,1311,1315,1318,1322,1325,1331,1334,1340,1348,1352,1355,1361,1364,1368,1371,1382],[11,1205,1206],{},"OpenAI sends different agents for different jobs, and the robots.txt setting for each is independent of the others. OpenAI's own example: allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot to keep your content out of training.",[15,1208,1210],{"id":1209},"openais-crawlers-at-a-glance","OpenAI's crawlers at a glance",[20,1212,1213,1226],{},[23,1214,1215],{},[26,1216,1217,1219,1221,1224],{},[29,1218,321],{},[29,1220,34],{},[29,1222,1223],{},"What robots.txt does",[29,1225,329],{},[39,1227,1228,1245,1262],{},[26,1229,1230,1234,1237,1240],{},[44,1231,1232],{},[47,1233,362],{},[44,1235,1236],{},"Finds websites to show in ChatGPT's search features",[44,1238,1239],{},"A Disallow keeps your pages out of ChatGPT search answers; they can still appear as navigational links",[44,1241,1242],{},[86,1243,374],{"href":372,"rel":1244},[90],[26,1246,1247,1251,1254,1257],{},[44,1248,1249],{},[47,1250,341],{},[44,1252,1253],{},"Crawls content that may be used to train OpenAI's generative AI foundation models",[44,1255,1256],{},"A Disallow tells OpenAI not to use your content for training",[44,1258,1259],{},[86,1260,353],{"href":351,"rel":1261},[90],[26,1263,1264,1268,1271,1274],{},[44,1265,1266],{},[47,1267,383],{},[44,1269,1270],{},"Visits a page for certain user actions in ChatGPT and custom GPTs, such as answering a question",[44,1272,1273],{},"Rules may not apply, because a user started the action",[44,1275,1276],{},[86,1277,395],{"href":393,"rel":1278},[90],[11,1280,1281,1282,1285,1286,92],{},"A fourth agent, ",[47,1283,1284],{},"OAI-AdsBot",", visits only the landing pages of ads submitted to ChatGPT to check that they follow OpenAI's ad policies. OpenAI says the data it collects isn't used to train foundation models. Its IP list is ",[86,1287,1290],{"href":1288,"rel":1289},"https:\u002F\u002Fopenai.com\u002Fadsbot.json",[90],"adsbot.json",[15,1292,1294],{"id":1293},"chatgpt-user-the-visit-behind-an-answer","ChatGPT-User: the visit behind an answer",[11,1296,1297],{},"When someone asks ChatGPT something that needs a web page, ChatGPT may visit it with the ChatGPT-User agent. OpenAI says that because a user initiated the visit, robots.txt rules may not apply. It also says ChatGPT-User isn't used to crawl the web automatically or to decide whether content appears in ChatGPT search.",[11,1299,1300],{},"So a ChatGPT-User request in your logs means a person's conversation needed that page at that moment. The request shows that ChatGPT read the page, not whether the answer quoted or linked it.",[15,1302,1304],{"id":1303},"gptbot-training-data","GPTBot: training data",[11,1306,1307],{},"GPTBot crawls content that may be used to train OpenAI's generative AI foundation models. Disallowing it tells OpenAI your content shouldn't be used for training. It doesn't decide whether you appear in ChatGPT search; that's OAI-SearchBot's job. If you allow both, OpenAI may use the results of one crawl for both purposes to avoid crawling twice.",[11,1309,1310],{},"OpenAI applies the same GPTBot opt-out to page content it gets through people's use of its ChatGPT Atlas browser. Even when an Atlas user opts in to training, pages that disallow GPTBot aren't trained on.",[15,1312,1314],{"id":1313},"oai-searchbot-chatgpt-search","OAI-SearchBot: ChatGPT search",[11,1316,1317],{},"OAI-SearchBot finds the pages ChatGPT's search features show and link to. For your content to be included in summaries and snippets in ChatGPT, OpenAI says not to block it. Any public website can appear in ChatGPT search. After you change robots.txt, it can take about 24 hours for ChatGPT search to adjust.",[15,1319,1321],{"id":1320},"robotstxt-rules-for-openais-crawlers","robots.txt rules for OpenAI's crawlers",[11,1323,1324],{},"Stay in ChatGPT search but opt out of training:",[115,1326,1329],{"className":1327,"code":1328,"language":120},[118],"User-agent: OAI-SearchBot\nAllow: \u002F\n\nUser-agent: GPTBot\nDisallow: \u002F\n",[47,1330,1328],{"__ignoreMap":123},[11,1332,1333],{},"Opt out of ChatGPT search as well:",[115,1335,1338],{"className":1336,"code":1337,"language":120},[118],"User-agent: OAI-SearchBot\nDisallow: \u002F\n\nUser-agent: GPTBot\nDisallow: \u002F\n",[47,1339,1337],{"__ignoreMap":123},[11,1341,144,1342,148,1344,152,1346,156],{},[47,1343,147],{},[47,1345,151],{},[47,1347,155],{},[15,1349,1351],{"id":1350},"user-agent-strings","User agent strings",[11,1353,1354],{},"OpenAI publishes these strings. It calls the GPTBot and OAI-SearchBot strings examples, so match on the agent's name rather than the whole string.",[115,1356,1359],{"className":1357,"code":1358,"language":120},[118],"OAI-SearchBot (example):\nMozilla\u002F5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit\u002F537.36 (KHTML, like Gecko) Chrome\u002F131.0.0.0 Safari\u002F537.36; compatible; OAI-SearchBot\u002F1.4; +https:\u002F\u002Fopenai.com\u002Fsearchbot\n\nGPTBot (example):\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko); compatible; GPTBot\u002F1.4; +https:\u002F\u002Fopenai.com\u002Fgptbot\n\nChatGPT-User:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko); compatible; ChatGPT-User\u002F1.0; +https:\u002F\u002Fopenai.com\u002Fbot\n",[47,1360,1358],{"__ignoreMap":123},[11,1362,1363],{},"Any script can send these strings. To confirm a request came from OpenAI, check its source IP address against the list for that agent in the table above.",[15,1365,1367],{"id":1366},"where-to-see-openais-crawlers-and-the-visitors-chatgpt-sends","Where to see OpenAI's crawlers and the visitors ChatGPT sends",[11,1369,1370],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so OpenAI's crawlers won't appear there. Your server and CDN logs record every request with its user agent and IP address.",[11,1372,1373,1374,1377,1378,92],{},"People who click a link in ChatGPT's search results are different: ChatGPT adds ",[47,1375,1376],{},"utm_source=chatgpt.com"," to those links, so the visits show up in your analytics under that source. To find them in GA4, see ",[86,1379,1381],{"href":1380},"\u002Fai-visibility\u002Ftrack-ai-traffic","how to track ChatGPT and AI traffic",[11,1383,1384,1386],{},[86,1385,192],{"href":191}," puts both sides in one report when your server reports page requests to OneLence: the pages GPTBot and OAI-SearchBot read, the pages ChatGPT fetched while answering people, the visitors ChatGPT sent to each page and the ones that converted.",{"title":123,"searchDepth":195,"depth":195,"links":1388},[1389,1390,1391,1392,1393,1394,1395],{"id":1209,"depth":195,"text":1210},{"id":1293,"depth":195,"text":1294},{"id":1303,"depth":195,"text":1304},{"id":1313,"depth":195,"text":1314},{"id":1320,"depth":195,"text":1321},{"id":1350,"depth":195,"text":1351},{"id":1366,"depth":195,"text":1367},"OpenAI runs separate agents for ChatGPT search (OAI-SearchBot), model training (GPTBot) and the visits people's requests trigger (ChatGPT-User). What each does, its user agent and IP list, and the robots.txt rules that control it.",[1398,1401,1404,1407,1410,1413,1416],{"question":1399,"answer":1400},"What is ChatGPT-User?","ChatGPT-User is the user agent ChatGPT uses when a person's request needs a web page, for example when they ask a question in ChatGPT or use a custom GPT. OpenAI says it isn't used to crawl the web automatically or to decide what appears in ChatGPT search.",{"question":1402,"answer":1403},"Does ChatGPT-User follow robots.txt?","Not necessarily. OpenAI says that because these actions are initiated by a user, robots.txt rules may not apply. GPTBot and OAI-SearchBot, which crawl automatically, are the agents robots.txt controls.",{"question":1405,"answer":1406},"What's the difference between GPTBot and OAI-SearchBot?","GPTBot crawls content that may be used to train OpenAI's generative AI foundation models. OAI-SearchBot finds websites to show in ChatGPT's search features. OpenAI treats the two settings independently, so you can allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot.",{"question":1408,"answer":1409},"Does blocking GPTBot remove my site from ChatGPT?","No. Blocking GPTBot tells OpenAI not to use your content for training. Whether your pages appear in ChatGPT search depends on OAI-SearchBot, and ChatGPT can still visit a page as ChatGPT-User when someone's request needs it.",{"question":1411,"answer":1412},"How do I get my site into ChatGPT search?","Don't block OAI-SearchBot. OpenAI says any public website can appear in ChatGPT search, and that sites opted out of OAI-SearchBot won't be shown in ChatGPT search answers, though they can still appear as navigational links. Changes to robots.txt can take about 24 hours to reach ChatGPT search.",{"question":1414,"answer":1415},"How do I track visitors from ChatGPT?","ChatGPT adds utm_source=chatgpt.com to the links people click in its search results, so those visits show up under that source in your analytics. ChatGPT-User requests aren't visitors: they are automated fetches, which belong in your server logs, not in your visitor numbers.",{"question":1417,"answer":1418},"How can I tell whether a request really came from OpenAI?","Check the source IP address against the list OpenAI publishes for that agent: openai.com\u002Fsearchbot.json for OAI-SearchBot, openai.com\u002Fgptbot.json for GPTBot and openai.com\u002Fchatgpt-user.json for ChatGPT-User. Anyone can copy a user agent string.",{},"ChatGPT-User, GPTBot and OAI-SearchBot",{"title":1422,"description":1423},"ChatGPT-User vs GPTBot vs OAI-SearchBot: OpenAI Crawlers Explained","What ChatGPT-User, GPTBot and OAI-SearchBot do, their user agents and IP lists, which follow robots.txt, and how to opt out of training while staying in ChatGPT search.",{"loc":834},"ai-crawlers\u002Fchatgpt-user","ChatGPT-User visits a page when someone's request in ChatGPT needs it. GPTBot collects content that may be used to train OpenAI's models, and OAI-SearchBot finds pages for ChatGPT search. Each has its own robots.txt setting.","R97y2LjvYj6iLK83zzDoFYN7PMA-RKsvUZAbk4vYNBY",{"id":1429,"title":1430,"body":1431,"checked":202,"description":1742,"documented":204,"extension":205,"faq":1743,"meta":1762,"name":885,"navigation":204,"order":1763,"path":884,"seo":1764,"sitemap":1767,"stem":1768,"tagline":1769,"__hash__":1770},"aiCrawlers\u002Fai-crawlers\u002Fllms-txt.md","llms.txt: what it is, who reads it and how to write one",{"type":8,"value":1432,"toc":1732},[1433,1437,1446,1449,1453,1456,1483,1490,1500,1504,1513,1519,1523,1538,1549,1563,1574,1578,1608,1611,1615,1670,1673,1677,1714,1718,1724],[15,1434,1436],{"id":1435},"what-llmstxt-is","What llms.txt is",[11,1438,1439,1440,1445],{},"Web pages are built for people, with navigation, scripts and other page furniture around the text. The ",[86,1441,1444],{"href":1442,"rel":1443},"https:\u002F\u002Fllmstxt.org\u002F",[90],"llms.txt proposal"," gives agents a shortcut: one markdown file that says what a site is, with links to the pages worth reading. In the proposal's words, agents view or search the llms.txt to find what they need, then follow the relevant links.",[11,1447,1448],{},"Jeremy Howard published the proposal on September 3, 2024. Version 2, from August 2026, adds standard ways for agents to find the file and the markdown versions of pages.",[15,1450,1452],{"id":1451},"the-format","The format",[11,1454,1455],{},"An llms.txt file is markdown with these parts, in this order:",[1457,1458,1459,1465,1471,1477],"ol",{},[825,1460,1461,1464],{},[913,1462,1463],{},"An H1 with the name of the site or project."," It's the only required part.",[825,1466,1467,1470],{},[913,1468,1469],{},"A blockquote with a short summary",", holding the key information needed to understand the rest of the file.",[825,1472,1473,1476],{},[913,1474,1475],{},"Any number of paragraphs or lists",", without headings, with more detail about the site and how to read the linked files.",[825,1478,1479,1482],{},[913,1480,1481],{},"Any number of sections under H2 headings",", each a list of links. Every item is a markdown link, optionally followed by a colon and a note about the page.",[11,1484,1485,1486,1489],{},"A section called ",[47,1487,1488],{},"Optional"," holds secondary links an agent can skip when it needs a shorter context.",[11,1491,1492,1493,1495,1496,1499],{},"To draft a file in this format from your own pages, use the free ",[86,1494,903],{"href":902},": it reads your home page, sitemaps and main pages, and you edit the result before you download it. To validate a live file against this format, enter your domain in the free ",[86,1497,1498],{"href":892},"robots.txt and llms.txt checker",". It flags a missing H1, sections without a link list and malformed links.",[15,1501,1503],{"id":1502},"an-example","An example",[11,1505,1506,1507,1512],{},"This is the start of ",[86,1508,1511],{"href":1509,"rel":1510},"https:\u002F\u002Fonelence.com\u002Fllms.txt",[90],"onelence.com\u002Fllms.txt",", shortened:",[115,1514,1517],{"className":1515,"code":1516,"language":120},[118],"# OneLence\n\n> OneLence is the growth operating system for founders and marketing teams. It watches the whole growth engine (website journeys and conversions, ad platforms, search, AI assistants and affiliate partners) and says what deserves attention, what to do about it and how sure it is, even when attribution is incomplete.\n\n## Connect OneLence to AI assistants\n\n- [OneLence MCP server](https:\u002F\u002Fonelence.com\u002Fintegrations\u002Fmcp): https:\u002F\u002Fmcp.onelence.com (Streamable HTTP, OAuth), available as a Claude connector\n- [Use OneLence in ChatGPT](https:\u002F\u002Fonelence.com\u002Fintegrations\u002Fchatgpt): turn on developer mode, then create an app with the server URL and OAuth.\n\n## Optional\n\n- [Blog](https:\u002F\u002Fonelence.com\u002Fblog)\n",[47,1518,1516],{"__ignoreMap":123},[15,1520,1522],{"id":1521},"markdown-versions-of-pages","Markdown versions of pages",[11,1524,1525,1526,1529,1530,1533,1534,1537],{},"The proposal also suggests that pages agents might need offer a clean markdown version at the same URL, with ",[47,1527,1528],{},".md"," appended (",[47,1531,1532],{},"page.html.md",") or with the extension replaced (",[47,1535,1536],{},"page.md","). Version 2 accepts both forms.",[11,1539,1540,1541,1544,1545,1548],{},"To help agents find these files, version 2 recommends two standard link relations, as HTML ",[47,1542,1543],{},"\u003Clink>"," elements or an HTTP ",[47,1546,1547],{},"Link"," header:",[822,1550,1551,1557],{},[825,1552,1553,1556],{},[47,1554,1555],{},"rel=\"alternate\" type=\"text\u002Fmarkdown\""," points to the markdown version of a page.",[825,1558,1559,1562],{},[47,1560,1561],{},"rel=\"describedby\""," points to the llms.txt file that covers the page.",[11,1564,1565,1566,1569,1570,1573],{},"An llms.txt file covers every page under its path, so ",[47,1567,1568],{},"\u002Fdocs\u002Fllms.txt"," covers everything in ",[47,1571,1572],{},"\u002Fdocs\u002F",", and a more specific file takes precedence.",[15,1575,1577],{"id":1576},"who-reads-llmstxt","Who reads llms.txt",[822,1579,1580,1586,1598],{},[825,1581,1582,1585],{},[913,1583,1584],{},"Google Search doesn't."," Google says you don't need machine-readable files, AI text files or markdown to appear in Google Search, including its generative AI features, because Google Search doesn't use them. It also says creating llms.txt files for other services is fine and will neither harm nor help your visibility or rankings in Google Search.",[825,1587,1588,1591,1592,180,1595,1597],{},[913,1589,1590],{},"The AI companies publish one, but don't say they read yours."," OpenAI, Anthropic and Perplexity each publish an llms.txt for their developer docs. Their crawler documentation doesn't say ",[86,1593,1594],{"href":834},"GPTBot, ChatGPT-User",[86,1596,49],{"href":227}," or PerplexityBot read llms.txt on other sites.",[825,1599,1600,1603,1604,1607],{},[913,1601,1602],{},"Lighthouse checks it."," Chrome's Lighthouse has an llms.txt audit among its agentic browsing audits. It fails only when your server returns an error for ",[47,1605,1606],{},"\u002Fllms.txt",". A missing file is marked not applicable, and the audit's documentation calls the file optional for now.",[11,1609,1610],{},"The only way to know whether agents read yours is to look at who requests it.",[15,1612,1614],{"id":1613},"llmstxt-vs-robotstxt-vs-sitemapxml","llms.txt vs robots.txt vs sitemap.xml",[20,1616,1617,1630],{},[23,1618,1619],{},[26,1620,1621,1624,1627],{},[29,1622,1623],{},"File",[29,1625,1626],{},"What it's for",[29,1628,1629],{},"Who it's written for",[39,1631,1632,1645,1658],{},[26,1633,1634,1639,1642],{},[44,1635,1636],{},[47,1637,1638],{},"robots.txt",[44,1640,1641],{},"Says which parts of a site crawlers may access",[44,1643,1644],{},"Every crawler, before it fetches pages",[26,1646,1647,1652,1655],{},[44,1648,1649],{},[47,1650,1651],{},"sitemap.xml",[44,1653,1654],{},"Lists every indexable page",[44,1656,1657],{},"Search engines building an index",[26,1659,1660,1664,1667],{},[44,1661,1662],{},[47,1663,885],{},[44,1665,1666],{},"Summarizes the site and links to the pages worth reading",[44,1668,1669],{},"An agent that needs information while helping someone",[11,1671,1672],{},"llms.txt doesn't replace either file, and it doesn't grant or block access.",[15,1674,1676],{"id":1675},"how-to-write-a-useful-llmstxt","How to write a useful llms.txt",[822,1678,1679,1685,1691,1699,1705],{},[825,1680,1681,1684],{},[913,1682,1683],{},"Put the facts you want repeated in the summary."," What you do, for whom, and anything an assistant tends to get wrong, such as pricing or which platforms you support.",[825,1686,1687,1690],{},[913,1688,1689],{},"Link the pages that answer real questions",": product, pricing, integrations, docs and comparisons, each with a one-line note on what the page answers.",[825,1692,1693,1698],{},[913,1694,1695,1696,92],{},"Move the rest under ",[47,1697,1488],{}," Blog posts and secondary pages go there.",[825,1700,1701,1704],{},[913,1702,1703],{},"Keep it current."," A summary with last year's prices does more harm than no file.",[825,1706,1707,1710,1711,1713],{},[913,1708,1709],{},"Check that it's fetchable"," at ",[47,1712,1606],{}," with a 200 response, and that robots.txt doesn't block it.",[15,1715,1717],{"id":1716},"seeing-who-fetches-your-llmstxt","Seeing who fetches your llms.txt",[11,1719,1720,1721,1723],{},"Your server and CDN logs record each request for ",[47,1722,1606],{}," with its user agent and IP address, which is enough to tell a browser from a crawler. Google Analytics won't show these requests, because it excludes known bots and only measures pages that load its tag.",[11,1725,1726,1728,1729,1731],{},[86,1727,192],{"href":191}," shows the pages AI crawlers request when your server reports those requests to OneLence. Report ",[47,1730,1606],{}," along with your pages and you can see which AI crawlers fetch it, and whether the pages it links to get read and visited.",{"title":123,"searchDepth":195,"depth":195,"links":1733},[1734,1735,1736,1737,1738,1739,1740,1741],{"id":1435,"depth":195,"text":1436},{"id":1451,"depth":195,"text":1452},{"id":1502,"depth":195,"text":1503},{"id":1521,"depth":195,"text":1522},{"id":1576,"depth":195,"text":1577},{"id":1613,"depth":195,"text":1614},{"id":1675,"depth":195,"text":1676},{"id":1716,"depth":195,"text":1717},"llms.txt is a markdown file that tells AI agents what a site is about and which pages to read. What the proposal says, what Google and the AI companies say about it, an example file, and how to see who fetches yours.",[1744,1747,1750,1753,1756,1759],{"question":1745,"answer":1746},"What is llms.txt?","llms.txt is a markdown file, usually at \u002Fllms.txt, that gives AI agents a short summary of a site and links to the pages with more detail. Jeremy Howard proposed it at llmstxt.org on September 3, 2024, and published version 2 of the proposal in August 2026.",{"question":1748,"answer":1749},"Does Google use llms.txt?","No. Google says Google Search, including its generative AI features, doesn't use llms.txt files, and that having one will neither harm nor help your site's visibility or rankings in Google Search.",{"question":1751,"answer":1752},"Do ChatGPT, Claude and Perplexity read llms.txt?","Their crawler documentation doesn't say so. OpenAI, Anthropic and Perplexity publish llms.txt files for their own developer docs, but none of them says GPTBot, ClaudeBot, PerplexityBot or their other agents read llms.txt on other sites. Your server logs show whether any agent fetches yours.",{"question":1754,"answer":1755},"Is llms.txt the same as robots.txt?","No. robots.txt tells crawlers which parts of a site they may access. llms.txt suggests what's worth reading, for an agent that needs information about a topic while helping someone. It doesn't allow or block anything.",{"question":1757,"answer":1758},"Where do I put llms.txt?","At the root of your site, as \u002Fllms.txt. The proposal also allows files at a subpath, such as \u002Fdocs\u002Fllms.txt, which then cover the pages under that path.",{"question":1760,"answer":1761},"Should I create an llms.txt file?","It takes little time, and Google says it won't hurt your rankings. Write one if you want agents that look for it to find an accurate summary and your most useful pages, but don't expect it to change your visibility in Google, ChatGPT or Claude by itself.",{},3,{"title":1765,"description":1766},"What Is llms.txt? Format, Example and Who Actually Reads It","llms.txt explained: the format from llmstxt.org, an example file, what Google, OpenAI and Anthropic say about it, the Lighthouse check, and how to see which agents fetch yours.",{"loc":884},"ai-crawlers\u002Fllms-txt","llms.txt is a markdown file at the root of a site that tells AI agents what the site is and which pages are worth reading. It's a proposal, not a standard: Google Search ignores it, and OpenAI, Anthropic and Perplexity don't say their crawlers use it.","7icvVoCLF5l4PvmK5cKsQ55HXUmDAlvDA2I3o3DuSmE",{"id":1772,"title":1773,"body":1774,"checked":1949,"description":1950,"documented":204,"extension":205,"faq":1951,"meta":1970,"name":842,"navigation":204,"order":1971,"path":841,"seo":1972,"sitemap":1975,"stem":1976,"tagline":1977,"__hash__":1978},"aiCrawlers\u002Fai-crawlers\u002Fperplexitybot.md","PerplexityBot vs Perplexity-User: Perplexity's crawlers explained",{"type":8,"value":1775,"toc":1938},[1776,1779,1783,1833,1836,1840,1843,1846,1850,1853,1856,1860,1863,1867,1870,1876,1879,1885,1893,1901,1905,1908,1910,1913,1919,1923,1926,1930,1933],[11,1777,1778],{},"Perplexity runs two agents, and each has its own robots.txt setting: PerplexityBot builds the search index, and Perplexity-User fetches pages for people's questions. Perplexity says each setting works independently and that its systems can take up to 24 hours to reflect a change.",[15,1780,1782],{"id":1781},"perplexitys-crawlers-at-a-glance","Perplexity's crawlers at a glance",[20,1784,1785,1797],{},[23,1786,1787],{},[26,1788,1789,1791,1793,1795],{},[29,1790,321],{},[29,1792,34],{},[29,1794,1223],{},[29,1796,329],{},[39,1798,1799,1816],{},[26,1800,1801,1805,1808,1811],{},[44,1802,1803],{},[47,1804,459],{},[44,1806,1807],{},"Surfaces and links websites in Perplexity's search results. Not used to crawl content for AI foundation models",[44,1809,1810],{},"A Disallow stops indexing of the page's full or partial text. The domain, headline and a brief factual summary may still be indexed",[44,1812,1813],{},[86,1814,471],{"href":469,"rel":1815},[90],[26,1817,1818,1822,1825,1828],{},[44,1819,1820],{},[47,1821,480],{},[44,1823,1824],{},"Visits a page when someone's question in Perplexity needs it. Not used for web crawling or training",[44,1826,1827],{},"Generally ignored, because a user requested the fetch",[44,1829,1830],{},[86,1831,492],{"href":490,"rel":1832},[90],[11,1834,1835],{},"Perplexity also says it works with third-party crawlers to help build its search index, and that its agreements require them to respect robots.txt. It doesn't name them or their user agents, so you can't write rules for them.",[15,1837,1839],{"id":1838},"perplexitybot-perplexitys-search-index","PerplexityBot: Perplexity's search index",[11,1841,1842],{},"PerplexityBot finds the pages that Perplexity's search results show and link to. Perplexity says it only crawls content in compliance with robots.txt, and it recommends allowing PerplexityBot, and requests from its published IP ranges, if you want your site to appear in Perplexity's search results.",[11,1844,1845],{},"Blocking it doesn't make a page disappear entirely. Perplexity says it won't index the full or partial text of a page that disallows PerplexityBot, but it may still index the domain, the headline and a brief factual summary.",[15,1847,1849],{"id":1848},"perplexity-user-the-visit-behind-an-answer","Perplexity-User: the visit behind an answer",[11,1851,1852],{},"When someone asks Perplexity a question, it might visit a web page with the Perplexity-User agent to help give an accurate answer. Perplexity says that since a user requested the fetch, Perplexity-User generally ignores robots.txt rules. Keeping it out takes a firewall rule, which also keeps your pages out of the answers people ask for.",[11,1854,1855],{},"A Perplexity-User request in your logs means someone's question needed that page at that moment. It shows that Perplexity read the page, not whether the answer quoted or linked it.",[15,1857,1859],{"id":1858},"does-perplexity-train-on-your-content","Does Perplexity train on your content?",[11,1861,1862],{},"Perplexity says no. It doesn't build foundation models, so your content won't be used for AI model pre-training. There is no Perplexity training crawler, which is why the training opt-out rules for OpenAI, Anthropic and Google have no Perplexity line.",[15,1864,1866],{"id":1865},"robotstxt-rules-for-perplexity","robots.txt rules for Perplexity",[11,1868,1869],{},"Stay in Perplexity's search results. This is also what happens when your robots.txt doesn't mention PerplexityBot:",[115,1871,1874],{"className":1872,"code":1873,"language":120},[118],"User-agent: PerplexityBot\nAllow: \u002F\n",[47,1875,1873],{"__ignoreMap":123},[11,1877,1878],{},"Keep your pages' text out of Perplexity's index:",[115,1880,1883],{"className":1881,"code":1882,"language":120},[118],"User-agent: PerplexityBot\nDisallow: \u002F\n",[47,1884,1882],{"__ignoreMap":123},[11,1886,1887,1888,1890,1891,92],{},"A rule for ",[47,1889,480],{}," is generally ignored, so robots.txt can't keep it out. Perplexity doesn't document support for ",[47,1892,102],{},[11,1894,144,1895,148,1897,152,1899,156],{},[47,1896,147],{},[47,1898,151],{},[47,1900,155],{},[15,1902,1904],{"id":1903},"the-2025-dispute-with-cloudflare","The 2025 dispute with Cloudflare",[11,1906,1907],{},"In August 2025, Cloudflare reported that when PerplexityBot was blocked, Perplexity also crawled with an undeclared user agent that impersonated Chrome on macOS, from IP addresses not on Perplexity's lists. Cloudflare removed Perplexity from its verified bots. Perplexity replied the same day that Cloudflare had misattributed the traffic, which it said came from BrowserBase's automated browser service. The two accounts still disagree.",[15,1909,1351],{"id":1350},[11,1911,1912],{},"Perplexity publishes these strings:",[115,1914,1917],{"className":1915,"code":1916,"language":120},[118],"PerplexityBot:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko; compatible; PerplexityBot\u002F1.0; +https:\u002F\u002Fperplexity.ai\u002Fperplexitybot)\n\nPerplexity-User:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko; compatible; Perplexity-User\u002F1.0; +https:\u002F\u002Fperplexity.ai\u002Fperplexity-user)\n",[47,1918,1916],{"__ignoreMap":123},[15,1920,1922],{"id":1921},"how-to-verify-a-perplexity-request","How to verify a Perplexity request",[11,1924,1925],{},"Any script can send these strings. To confirm a request came from Perplexity, check its source IP address against the list for that agent in the table above. Perplexity recommends combining user agent matching with IP address verification when you write firewall rules for its bots. It doesn't document a reverse DNS check.",[15,1927,1929],{"id":1928},"where-to-see-perplexitys-crawlers-and-the-visitors-it-sends","Where to see Perplexity's crawlers and the visitors it sends",[11,1931,1932],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so PerplexityBot and Perplexity-User won't appear there. Your server and CDN logs record every request with its user agent and IP address. People who click a link in a Perplexity answer are different: they are visitors, and when their browser passes the referrer they show up in your analytics with perplexity.ai as the source.",[11,1934,1935,1937],{},[86,1936,192],{"href":191}," puts both sides in one report when your server reports page requests to OneLence. It marks PerplexityBot and Perplexity-User requests confirmed by Perplexity's IP lists as verified, and shows which pages PerplexityBot read, which pages Perplexity-User fetched while answering people, and the visitors Perplexity sent you and the ones that converted.",{"title":123,"searchDepth":195,"depth":195,"links":1939},[1940,1941,1942,1943,1944,1945,1946,1947,1948],{"id":1781,"depth":195,"text":1782},{"id":1838,"depth":195,"text":1839},{"id":1848,"depth":195,"text":1849},{"id":1858,"depth":195,"text":1859},{"id":1865,"depth":195,"text":1866},{"id":1903,"depth":195,"text":1904},{"id":1350,"depth":195,"text":1351},{"id":1921,"depth":195,"text":1922},{"id":1928,"depth":195,"text":1929},"2026-10-05","PerplexityBot builds the index behind Perplexity's search results, and Perplexity-User fetches pages while Perplexity answers someone. What each does, which one follows robots.txt, their IP lists, and what blocking them changes.",[1952,1955,1958,1961,1964,1967],{"question":1953,"answer":1954},"What is PerplexityBot?","PerplexityBot is Perplexity's search crawler. Perplexity says it is designed to surface and link websites in search results on Perplexity, and that it isn't used to crawl content for AI foundation models.",{"question":1956,"answer":1957},"What is the difference between PerplexityBot and Perplexity-User?","PerplexityBot crawls the web for Perplexity's search index and follows robots.txt. Perplexity-User visits a page when someone asks Perplexity a question that needs it. Perplexity says that because a user requested the fetch, Perplexity-User generally ignores robots.txt rules.",{"question":1959,"answer":1960},"Does Perplexity respect robots.txt?","PerplexityBot does: Perplexity says it only crawls content in compliance with robots.txt. Perplexity-User generally doesn't, because a person asked for the page. In 2025 Cloudflare reported undeclared crawling that it attributed to Perplexity, and Perplexity denied it.",{"question":1962,"answer":1963},"Will blocking PerplexityBot remove my site from Perplexity?","Not completely. Perplexity says it won't index the full or partial text of a page that disallows PerplexityBot, but it may still index the domain, the headline and a brief factual summary. Perplexity-User can also still visit when someone asks about your page, since it generally ignores robots.txt.",{"question":1965,"answer":1966},"Does Perplexity train AI models on my content?","Perplexity says no. It doesn't build foundation models, so your content won't be used for AI model pre-training. That's why there is no Perplexity training crawler to block.",{"question":1968,"answer":1969},"How can I tell whether a request really came from Perplexity?","Check the source IP address against the list Perplexity publishes for that agent: perplexity.com\u002Fperplexitybot.json for PerplexityBot and perplexity.com\u002Fperplexity-user.json for Perplexity-User. Perplexity recommends combining user agent matching with IP address checks in firewall rules.",{},4,{"title":1973,"description":1974},"PerplexityBot vs Perplexity-User: What They Do and How to Block Them","What PerplexityBot and Perplexity-User do, their user agents and IP lists, which one follows robots.txt, and what blocking PerplexityBot does to your place in Perplexity.",{"loc":841},"ai-crawlers\u002Fperplexitybot","PerplexityBot finds and links pages for Perplexity's search results and follows robots.txt. Perplexity-User visits a page when someone's question needs it and generally ignores robots.txt. Perplexity says neither collects content to train AI models.","FfCLNJrOSr7GdoQuMo9qMhFwXRqPRDP-wvPKSPPHFhk",{"id":1980,"title":1981,"body":1982,"checked":1949,"description":2224,"documented":204,"extension":205,"faq":2225,"meta":2243,"name":502,"navigation":204,"order":2244,"path":303,"seo":2245,"sitemap":2248,"stem":2249,"tagline":2250,"__hash__":2251},"aiCrawlers\u002Fai-crawlers\u002Fgoogle-extended.md","Google-Extended: what it controls, and how to opt out of AI Overviews",{"type":8,"value":1983,"toc":2214},[1984,1987,1991,2078,2082,2085,2098,2101,2104,2108,2111,2114,2118,2121,2124,2135,2141,2145,2148,2154,2157,2163,2168,2176,2180,2183,2186,2190,2202,2206,2209],[11,1985,1986],{},"Google-Extended is the robots.txt name Google gives to one decision: whether the pages it crawls for other purposes may also be used for Gemini. It doesn't crawl anything itself, so it never shows up in your logs.",[15,1988,1990],{"id":1989},"googles-ai-controls-at-a-glance","Google's AI controls at a glance",[20,1992,1993,2006],{},[23,1994,1995],{},[26,1996,1997,2000,2003],{},[29,1998,1999],{},"Control",[29,2001,2002],{},"What it changes",[29,2004,2005],{},"What it doesn't change",[39,2007,2008,2022,2033,2045,2066],{},[26,2009,2010,2016,2019],{},[44,2011,2012,2013,2015],{},"Disallow ",[47,2014,502],{}," in robots.txt",[44,2017,2018],{},"Training of future Gemini models, grounding in Gemini Apps and Grounding with Google Search on Vertex AI, and training of the models behind Search's generative AI features",[44,2020,2021],{},"Inclusion and ranking in Google Search, which includes AI Overviews and AI Mode",[26,2023,2024,2027,2030],{},[44,2025,2026],{},"Search generative AI control in Search Console",[44,2028,2029],{},"Links to your site and your content in AI Overviews, AI Mode and generative AI features in Discover",[44,2031,2032],{},"The rest of Google Search",[26,2034,2035,2040,2043],{},[44,2036,2012,2037,2015],{},[47,2038,2039],{},"Googlebot",[44,2041,2042],{},"All of Google Search, AI features included, plus Google Images, Google Video and Google News",[44,2044],{},[26,2046,2047,2061,2064],{},[44,2048,2049,180,2052,180,2055,180,2058],{},[47,2050,2051],{},"nosnippet",[47,2053,2054],{},"data-nosnippet",[47,2056,2057],{},"max-snippet",[47,2059,2060],{},"noindex",[44,2062,2063],{},"What Search shows from the page",[44,2065],{},[26,2067,2068,2073,2076],{},[44,2069,2012,2070,2015],{},[47,2071,2072],{},"GoogleOther",[44,2074,2075],{},"No specific product, says Google",[44,2077],{},[15,2079,2081],{"id":2080},"what-google-extended-controls","What Google-Extended controls",[11,2083,2084],{},"Google describes Google-Extended as a standalone product token that sites can use to manage whether content Google crawls from them may be used for two things:",[822,2086,2087,2092],{},[825,2088,2089,2091],{},[913,2090,269],{}," future generations of the Gemini models that power Gemini Apps and the Vertex AI API for Gemini.",[825,2093,2094,2097],{},[913,2095,2096],{},"Grounding"," in Gemini Apps and Grounding with Google Search on Vertex AI, which means providing content from the Google Search index to the model at prompt time to make answers more factual and relevant.",[11,2099,2100],{},"Search Console's help adds one more: to limit training of the models used to generate responses in Search's generative AI features, use Google-Extended.",[11,2102,2103],{},"Google says Google-Extended doesn't affect a site's inclusion in Google Search and isn't used as a ranking signal. Blocking it isn't an SEO decision.",[15,2105,2107],{"id":2106},"ai-overviews-and-ai-mode-are-part-of-search","AI Overviews and AI Mode are part of Search",[11,2109,2110],{},"Google says AI is built into Search, which is why robots.txt rules for Googlebot are the control for how sites are crawled for Search. To be eligible as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to appear in Google Search with a snippet, and Google says there are no additional requirements or special optimizations.",[11,2112,2113],{},"So blocking Google-Extended doesn't keep your pages out of AI Overviews or AI Mode.",[15,2115,2117],{"id":2116},"the-search-generative-ai-control","The Search generative AI control",[11,2119,2120],{},"Search Console now has a separate setting for that. Google began testing it on June 3, 2026, with a subset of site owners in the UK, and says that as of August 31, 2026, it has rolled it out to all websites worldwide.",[11,2122,2123],{},"When you exclude your site with the control:",[822,2125,2126,2129,2132],{},[825,2127,2128],{},"Links to your site and your site's content won't appear in AI Overviews, AI Mode or generative AI features in Discover.",[825,2130,2131],{},"Content crawled from your site won't be used as an input to generate an AI response or preview in those features.",[825,2133,2134],{},"The rest of Google Search is unaffected. Google says the control isn't used as a ranking or inclusion signal for other parts of Search.",[11,2136,2137,2138,2140],{},"Google says content is excluded within 1 to 2 days after the control goes live, though caching can make some content take longer. To remove your content from Google Search completely, Google points to ",[47,2139,2060],{}," instead. Google doesn't say whether the control affects the Gemini app, so Google-Extended remains the documented setting for Gemini Apps.",[15,2142,2144],{"id":2143},"robotstxt-rules-for-google-extended","robots.txt rules for Google-Extended",[11,2146,2147],{},"Opt out of Gemini training and grounding while staying in Google Search:",[115,2149,2152],{"className":2150,"code":2151,"language":120},[118],"User-agent: Google-Extended\nDisallow: \u002F\n",[47,2153,2151],{"__ignoreMap":123},[11,2155,2156],{},"Opt out for part of your site, following Google's own example:",[115,2158,2161],{"className":2159,"code":2160,"language":120},[118],"User-agent: Google-Extended\nAllow: \u002Farchive\u002F1Q84\nDisallow: \u002Farchive\u002F\n",[47,2162,2160],{"__ignoreMap":123},[11,2164,2165,2166,92],{},"Google's crawlers always obey robots.txt rules when they crawl automatically. Google doesn't support ",[47,2167,102],{},[11,2169,144,2170,148,2172,152,2174,156],{},[47,2171,147],{},[47,2173,151],{},[47,2175,155],{},[15,2177,2179],{"id":2178},"googleother-and-google-agent","GoogleOther and Google-Agent",[11,2181,2182],{},"GoogleOther is a generic crawler that Google's product teams may use to fetch publicly accessible content, for example for one-off crawls for internal research and development. Google says robots.txt rules for GoogleOther don't affect any specific product.",[11,2184,2185],{},"Google-Agent is different: agents hosted on Google infrastructure use it to navigate the web and perform actions when a user asks. Because a user requested the fetch, Google says these fetchers generally ignore robots.txt rules.",[15,2187,2189],{"id":2188},"how-to-verify-googles-crawlers","How to verify Google's crawlers",[11,2191,2192,2193,2198,2199,92],{},"Google-Extended sends no requests, so there's nothing to verify. For the crawlers that do the crawling, Googlebot and GoogleOther among them, check the source IP address against Google's ",[86,2194,2197],{"href":2195,"rel":2196},"https:\u002F\u002Fdevelopers.google.com\u002Fstatic\u002Fcrawling\u002Fipranges\u002Fcommon-crawlers.json",[90],"common-crawlers.json",", or run a reverse DNS lookup: the hostname should be in googlebot.com, google.com or googleusercontent.com, and a forward lookup of that hostname should return the original IP address. Google-Agent uses the addresses in ",[86,2200,532],{"href":530,"rel":2201},[90],[15,2203,2205],{"id":2204},"where-to-see-googles-crawlers","Where to see Google's crawlers",[11,2207,2208],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, so Googlebot won't appear there. Your server and CDN logs record its requests. Google-Extended won't appear anywhere, since no request carries it.",[11,2210,2211,2213],{},[86,2212,192],{"href":191}," counts Googlebot as classic search rather than AI. It shows the visitors Gemini sends you and the pages Google's user-triggered agents, such as Google-Agent, fetch, confirmed against Google's IP lists, next to the same view for ChatGPT, Claude and Perplexity.",{"title":123,"searchDepth":195,"depth":195,"links":2215},[2216,2217,2218,2219,2220,2221,2222,2223],{"id":1989,"depth":195,"text":1990},{"id":2080,"depth":195,"text":2081},{"id":2106,"depth":195,"text":2107},{"id":2116,"depth":195,"text":2117},{"id":2143,"depth":195,"text":2144},{"id":2178,"depth":195,"text":2179},{"id":2188,"depth":195,"text":2189},{"id":2204,"depth":195,"text":2205},"Google-Extended is a robots.txt token, not a crawler. What it controls (Gemini training and grounding), what it doesn't (Google Search, AI Overviews and AI Mode), and the Search Console setting that opts a site out of AI Overviews and AI Mode.",[2226,2229,2232,2235,2237,2240],{"question":2227,"answer":2228},"What is Google-Extended?","A robots.txt product token that lets sites choose whether content Google crawls from them may be used to train future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It has no user agent of its own: Google crawls with its existing user agents.",{"question":2230,"answer":2231},"Why don't I see Google-Extended in my server logs?","Because it isn't a crawler. Google says Google-Extended has no separate HTTP request user agent string and that crawling is done with Google's existing user agents. The token only controls how Google may use what it crawls.",{"question":2233,"answer":2234},"Does blocking Google-Extended hurt my SEO?","No. Google says Google-Extended doesn't affect a site's inclusion in Google Search and isn't used as a ranking signal in Google Search.",{"question":1042,"answer":2236},"No. AI Overviews and AI Mode are part of Google Search, which Google-Extended doesn't affect. To opt out of them, use the Search generative AI control in Search Console, which Google rolled out to all websites worldwide as of August 31, 2026. It removes links to your site and your content from AI Overviews, AI Mode and generative AI features in Discover without affecting the rest of Search.",{"question":2238,"answer":2239},"How do I opt out of AI Overviews and AI Mode?","Use the Search generative AI control in Search Console. Google says content is excluded within 1 to 2 days after the control goes live, though caching can make some content take longer, and that the control isn't used as a ranking or inclusion signal for the rest of Search. To limit what Search shows from a page, Google also supports nosnippet, data-nosnippet, max-snippet and noindex.",{"question":2241,"answer":2242},"What is GoogleOther?","A generic Google crawler that product teams may use to fetch publicly accessible content, for example for one-off crawls for internal research and development. Google says robots.txt rules for GoogleOther don't affect any specific product.",{},5,{"title":2246,"description":2247},"Google-Extended Explained: Gemini Training, AI Overviews and Opt-Outs","What Google-Extended controls, why it never shows in your logs, why blocking it doesn't remove you from AI Overviews, and the Search Console setting that does.",{"loc":303},"ai-crawlers\u002Fgoogle-extended","Google-Extended is a robots.txt setting, not a crawler: it decides whether pages Google already crawls may train Gemini models and ground Gemini's answers. It doesn't affect Google Search, AI Overviews or AI Mode. Since August 31, 2026, Search Console has a separate setting for those.","It-fspyy8X5E_oY0FPE52LkJ76lkVWApBJ9YLMGVum4",{"id":2253,"title":2254,"body":2255,"checked":1949,"description":2471,"documented":204,"extension":205,"faq":2472,"meta":2488,"name":542,"navigation":204,"order":2489,"path":853,"seo":2490,"sitemap":2493,"stem":2494,"tagline":2495,"__hash__":2496},"aiCrawlers\u002Fai-crawlers\u002Famazonbot.md","Amazonbot: what Amazon's crawlers do and how to control them",{"type":8,"value":2256,"toc":2463},[2257,2260,2264,2334,2338,2341,2377,2381,2384,2390,2393,2399,2405,2411,2419,2421,2428,2434,2438,2445,2449,2458],[11,2258,2259],{},"Amazon documents three agents on one page and says each user agent setting is independent of the others. Amazon says changes can take about 24 hours to reach its systems, though the same page also says the crawlers may use a copy of your robots.txt cached within the last 30 days.",[15,2261,2263],{"id":2262},"amazons-crawlers-at-a-glance","Amazon's crawlers at a glance",[20,2265,2266,2278],{},[23,2267,2268],{},[26,2269,2270,2272,2274,2276],{},[29,2271,321],{},[29,2273,34],{},[29,2275,1223],{},[29,2277,329],{},[39,2279,2280,2298,2316],{},[26,2281,2282,2286,2289,2292],{},[44,2283,2284],{},[47,2285,542],{},[44,2287,2288],{},"Improves Amazon's products and services, and may be used to train Amazon AI models",[44,2290,2291],{},"A Disallow stops Amazonbot crawling those pages",[44,2293,2294],{},[86,2295,2297],{"href":553,"rel":2296},[90],"Amazonbot IP addresses",[26,2299,2300,2304,2307,2310],{},[44,2301,2302],{},[47,2303,564],{},[44,2305,2306],{},"Makes content eligible for search experiences such as Alexa. Not used for generative AI training",[44,2308,2309],{},"A Disallow stops it crawling. Amazon says allowing it is what makes your content eligible",[44,2311,2312],{},[86,2313,2315],{"href":574,"rel":2314},[90],"Amzn-SearchBot IP addresses",[26,2317,2318,2322,2325,2328],{},[44,2319,2320],{},[47,2321,585],{},[44,2323,2324],{},"Supports user actions, such as answering Alexa questions that need up-to-date information. Not used for generative AI training",[44,2326,2327],{},"May not follow every directive, because a user can start its actions",[44,2329,2330],{},[86,2331,2333],{"href":596,"rel":2332},[90],"Amzn-User IP addresses",[15,2335,2337],{"id":2336},"how-amazons-crawlers-treat-robotstxt","How Amazon's crawlers treat robots.txt",[11,2339,2340],{},"Amazon says its crawlers respect the Robots Exclusion Protocol, honoring the user-agent line and the allow and disallow directives. In detail:",[822,2342,2343,2346,2349,2358,2374],{},[825,2344,2345],{},"They fetch each host's robots.txt, or use a copy cached within the last 30 days. When the file can't be fetched, they behave as if it doesn't exist.",[825,2347,2348],{},"They honor the rules each host of a domain exposes, so every subdomain needs its own robots.txt.",[825,2350,2351,2352,2357],{},"They ",[913,2353,2354,2355],{},"don't support ",[47,2356,102],{},", and Amazon documents no other way to slow them down.",[825,2359,2360,2361,2364,2365,2368,2369,183,2371,92],{},"They respect ",[47,2362,2363],{},"rel=nofollow"," on links and the page-level robots meta tags ",[47,2366,2367],{},"noarchive"," (do not use the page for model training), ",[47,2370,2060],{},[47,2372,2373],{},"none",[825,2375,2376],{},"If robots.txt doesn't mention Amzn-SearchBot but allows other search bots, Amzn-SearchBot follows the rules given to those search bots.",[15,2378,2380],{"id":2379},"robotstxt-rules-for-amazons-crawlers","robots.txt rules for Amazon's crawlers",[11,2382,2383],{},"Opt out of Amazonbot, AI training included, while staying eligible for Amazon's search experiences such as Alexa:",[115,2385,2388],{"className":2386,"code":2387,"language":120},[118],"User-agent: Amazonbot\nDisallow: \u002F\n\nUser-agent: Amzn-SearchBot\nAllow: \u002F\n",[47,2389,2387],{"__ignoreMap":123},[11,2391,2392],{},"Opt out of all three:",[115,2394,2397],{"className":2395,"code":2396,"language":120},[118],"User-agent: Amazonbot\nDisallow: \u002F\n\nUser-agent: Amzn-SearchBot\nDisallow: \u002F\n\nUser-agent: Amzn-User\nDisallow: \u002F\n",[47,2398,2396],{"__ignoreMap":123},[11,2400,2401,2402,2404],{},"Amzn-User may still fetch a page when a user starts the action. To keep a single page out of training without blocking any crawler, Amazon honors the ",[47,2403,2367],{}," robots meta tag, which it defines as \"do not use the page for model training\":",[115,2406,2409],{"className":2407,"code":2408,"language":120},[118],"\u003Cmeta name=\"robots\" content=\"noarchive\">\n",[47,2410,2408],{"__ignoreMap":123},[11,2412,144,2413,148,2415,152,2417,156],{},[47,2414,147],{},[47,2416,151],{},[47,2418,155],{},[15,2420,1351],{"id":1350},[11,2422,2423,2424,2427],{},"Amazon publishes these strings. ",[47,2425,2426],{},"Chrome\u002FW.X.Y.Z"," stands for a Chrome version, so match on the agent's name rather than the whole string:",[115,2429,2432],{"className":2430,"code":2431,"language":120},[118],"Amazonbot:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko; compatible; Amazonbot\u002F0.1) Chrome\u002FW.X.Y.Z Safari\u002F537.36\n\nAmzn-SearchBot:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot\u002F0.1) Chrome\u002FW.X.Y.Z Safari\u002F537.36\n\nAmzn-User:\nMozilla\u002F5.0 AppleWebKit\u002F537.36 (KHTML, like Gecko; compatible; Amzn-User\u002F0.1) Chrome\u002FW.X.Y.Z Safari\u002F537.36\n",[47,2433,2431],{"__ignoreMap":123},[15,2435,2437],{"id":2436},"how-to-verify-an-amazon-request","How to verify an Amazon request",[11,2439,2440,2441,92],{},"Any script can send these strings. Amazon publishes the IP addresses each agent uses, linked in the table above, as lists of single addresses. Check a request's source IP address against the list for its agent. Amazon doesn't document a reverse DNS check, so don't rely on older guides that describe one. Publishers with questions can write to ",[86,2442,2444],{"href":2443},"mailto:amazonbot@amazon.com","amazonbot@amazon.com",[15,2446,2448],{"id":2447},"where-to-see-amazons-crawlers","Where to see Amazon's crawlers",[11,2450,2451,2452,180,2454,183,2456,92],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Amazonbot won't appear there. Your server and CDN logs record every request with its user agent and IP address. Search them for ",[47,2453,542],{},[47,2455,564],{},[47,2457,585],{},[11,2459,2460,2462],{},[86,2461,192],{"href":191}," does this for you when your server reports page requests to OneLence. It recognizes Amazonbot, Amzn-SearchBot and Amzn-User by their user agents and shows which pages each one requested, next to the same view for OpenAI, Anthropic, Perplexity and Google.",{"title":123,"searchDepth":195,"depth":195,"links":2464},[2465,2466,2467,2468,2469,2470],{"id":2262,"depth":195,"text":2263},{"id":2336,"depth":195,"text":2337},{"id":2379,"depth":195,"text":2380},{"id":1350,"depth":195,"text":1351},{"id":2436,"depth":195,"text":2437},{"id":2447,"depth":195,"text":2448},"Amazon runs three agents: Amazonbot, whose data may train Amazon AI models, Amzn-SearchBot for search experiences such as Alexa, and Amzn-User for live user requests. What each does, how robots.txt applies and how to verify them.",[2473,2476,2479,2482,2485],{"question":2474,"answer":2475},"What is Amazonbot?","Amazonbot is Amazon's web crawler. Amazon says it is used to improve Amazon's products and services, helps provide more accurate information to customers, and may be used to train Amazon AI models.",{"question":2477,"answer":2478},"Does Amazonbot respect robots.txt?","Yes. Amazon says Amazonbot, Amzn-SearchBot and Amzn-User respect the Robots Exclusion Protocol, honoring the user-agent line and the allow and disallow directives. They fetch each host's robots.txt or use a copy cached within the last 30 days. They don't support Crawl-delay, and Amzn-User may not follow every directive because a user can start its actions.",{"question":2480,"answer":2481},"Should I block Amazonbot?","Block it if you don't want your content used to train Amazon AI models. A rule for Amazonbot doesn't apply to Amzn-SearchBot, which makes your content eligible for search experiences such as Alexa, or to Amzn-User, which fetches pages for live requests. Amazon says each user agent setting is independent of the others.",{"question":2483,"answer":2484},"How do I slow Amazonbot down?","robots.txt can't do it: Amazon says its crawlers don't support the Crawl-delay directive, and it documents no other way to set a crawl rate. You can disallow the paths you don't want crawled, or write to Amazon at amazonbot@amazon.com.",{"question":2486,"answer":2487},"How can I tell whether a request really came from Amazon?","Check the source IP address against the address list Amazon publishes for that agent on its developer site: one each for Amazonbot, Amzn-SearchBot and Amzn-User. Amazon doesn't document a reverse DNS check, so don't rely on older guides that describe one.",{},6,{"title":2491,"description":2492},"Amazonbot: Amazon's Crawler, Amzn-SearchBot and How to Block Them","What Amazonbot, Amzn-SearchBot and Amzn-User do, their user agents and IP lists, how robots.txt applies (no Crawl-delay), and how to opt out of training but stay in Alexa.",{"loc":853},"ai-crawlers\u002Famazonbot","Amazonbot helps improve Amazon's products and may be used to train Amazon AI models. Amzn-SearchBot makes pages eligible for search experiences such as Alexa, and Amzn-User fetches pages for live requests. Each has its own robots.txt setting, and none supports Crawl-delay.","1XXTugLg2K60jjeQR8O7Y9YtdkKnLtBMmt9OKIaMFGI",{"id":2498,"title":2499,"body":2500,"checked":1949,"description":2702,"documented":204,"extension":205,"faq":2703,"meta":2722,"name":860,"navigation":204,"order":2723,"path":859,"seo":2724,"sitemap":2727,"stem":2728,"tagline":2729,"__hash__":2730},"aiCrawlers\u002Fai-crawlers\u002Fapplebot.md","Applebot and Applebot-Extended: Apple's crawler and its AI training opt-out",{"type":8,"value":2501,"toc":2692},[2502,2505,2509,2550,2556,2560,2563,2566,2570,2573,2577,2600,2604,2607,2613,2616,2622,2631,2639,2641,2648,2654,2657,2661,2668,2674,2680,2684,2687],[11,2503,2504],{},"Apple runs one crawler, Applebot, and one extra robots.txt setting, Applebot-Extended, that controls how Applebot's data may be used for AI.",[15,2506,2508],{"id":2507},"applebot-and-applebot-extended-at-a-glance","Applebot and Applebot-Extended at a glance",[20,2510,2511,2524],{},[23,2512,2513],{},[26,2514,2515,2518,2521],{},[29,2516,2517],{},"Name",[29,2519,2520],{},"What it is",[29,2522,2523],{},"What a Disallow does",[39,2525,2526,2538],{},[26,2527,2528,2532,2535],{},[44,2529,2530],{},[47,2531,608],{},[44,2533,2534],{},"Apple's crawler. Its data powers search in Spotlight, Siri and Safari, and may help train Apple foundation models",[44,2536,2537],{},"Stops Applebot crawling those pages, for search and training alike",[26,2539,2540,2544,2547],{},[44,2541,2542],{},[47,2543,629],{},[44,2545,2546],{},"A robots.txt setting, not a crawler. It decides whether Applebot's data may train Apple's foundation models",[44,2548,2549],{},"Opts those pages out of training. They can still appear in search results",[11,2551,2552,2553,92],{},"Apple publishes Applebot's IP ranges in ",[86,2554,620],{"href":618,"rel":2555},[90],[15,2557,2559],{"id":2558},"what-applebot-does","What Applebot does",[11,2561,2562],{},"Apple says the data Applebot crawls powers features such as the search technology in Spotlight, Siri and Safari. The same data may also help train the Apple foundation models behind generative AI features across Apple products, including Apple Intelligence, Services and Developer Tools.",[11,2564,2565],{},"Applebot may render pages in a browser. Apple says that if robots.txt blocks your JavaScript, CSS or other resources, Applebot may not be able to render the content properly.",[15,2567,2569],{"id":2568},"applebot-extended-the-training-opt-out","Applebot-Extended: the training opt-out",[11,2571,2572],{},"Applebot-Extended doesn't crawl webpages. Apple says it is only used to determine how to use the data Applebot crawls, and that disallowing it opts your content out of being used to train Apple's general purpose foundation models. Pages that disallow Applebot-Extended can still be included in search results.",[15,2574,2576],{"id":2575},"how-applebot-treats-robotstxt","How Applebot treats robots.txt",[822,2578,2579,2582,2585,2593],{},[825,2580,2581],{},"Applebot respects standard robots.txt directives that target it.",[825,2583,2584],{},"If your robots.txt doesn't mention Applebot but mentions Googlebot, Applebot follows the Googlebot rules. A site that blocks Googlebot and never names Applebot blocks Applebot too.",[825,2586,2587,2588,92],{},"Applebot ",[913,2589,2590,2591],{},"doesn't follow ",[47,2592,102],{},[825,2594,2595,2596,2599],{},"It supports robots meta tags in HTML and indexing directives in the ",[47,2597,2598],{},"X-Robots-Tag"," HTTP header.",[15,2601,2603],{"id":2602},"robotstxt-rules-for-apple","robots.txt rules for Apple",[11,2605,2606],{},"Stay in search in Siri, Spotlight and Safari, but opt out of training Apple's foundation models:",[115,2608,2611],{"className":2609,"code":2610,"language":120},[118],"User-agent: Applebot-Extended\nDisallow: \u002F\n",[47,2612,2610],{"__ignoreMap":123},[11,2614,2615],{},"Opt out of training for one folder only, as in Apple's own example:",[115,2617,2620],{"className":2618,"code":2619,"language":120},[118],"User-agent: Applebot-Extended\nDisallow: \u002Fprivate\u002F\n",[47,2621,2619],{"__ignoreMap":123},[11,2623,2624,2625,2627,2628,92],{},"To keep a page in search but out of the context AI models use when they generate output in Apple products, Apple names two signals: it won't use content tagged ",[47,2626,2051],{}," that way, nor pages marked ",[47,2629,2630],{},"isAccessibleForFree: false",[11,2632,144,2633,148,2635,152,2637,156],{},[47,2634,147],{},[47,2636,151],{},[47,2638,155],{},[15,2640,1351],{"id":1350},[11,2642,2643,2644,2647],{},"Apple gives these as examples for desktop and mobile. The stable part is ",[47,2645,2646],{},"Applebot\u002F0.1; +http:\u002F\u002Fwww.apple.com\u002Fgo\u002Fapplebot",", so match on that:",[115,2649,2652],{"className":2650,"code":2651,"language":120},[118],"Desktop (example):\nMozilla\u002F5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit\u002F605.1.15 (KHTML, like Gecko) Version\u002F17.4 Safari\u002F605.1.15 (Applebot\u002F0.1; +http:\u002F\u002Fwww.apple.com\u002Fgo\u002Fapplebot)\n\nMobile (example):\nMozilla\u002F5.0 (iPhone; CPU iPhone OS 17_4_1 like Mac OS X) AppleWebKit\u002F605.1.15 (KHTML, like Gecko) Version\u002F17.4.1 Mobile\u002F15E148 Safari\u002F604.1 (Applebot\u002F0.1; +http:\u002F\u002Fwww.apple.com\u002Fgo\u002Fapplebot)\n",[47,2653,2651],{"__ignoreMap":123},[11,2655,2656],{},"Applebot-Extended has no user agent string, because it never sends a request.",[15,2658,2660],{"id":2659},"how-to-verify-applebot","How to verify Applebot",[11,2662,2663,2664,2667],{},"Apple says Applebot traffic is generally identified by reverse DNS in the ",[47,2665,2666],{},"applebot.apple.com"," domain. Apple's example:",[115,2669,2672],{"className":2670,"code":2671,"language":120},[118],"host 17.58.101.179\n# points to 17-58-101-179.applebot.apple.com\n\nhost 17-58-101-179.applebot.apple.com\n# should return 17.58.101.179\n",[47,2673,2671],{"__ignoreMap":123},[11,2675,2676,2677,92],{},"You can also match the IP address against the CIDR prefixes in ",[86,2678,620],{"href":618,"rel":2679},[90],[15,2681,2683],{"id":2682},"where-to-see-applebot","Where to see Applebot",[11,2685,2686],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Applebot won't appear there. Your server and CDN logs record its requests.",[11,2688,2689,2691],{},[86,2690,192],{"href":191}," counts Applebot as a search crawler, like Googlebot and Bingbot, rather than as AI, and confirms its requests against Apple's IP list. Because Applebot can render pages, the OneLence tracking tag can record it even without server-side tracking.",{"title":123,"searchDepth":195,"depth":195,"links":2693},[2694,2695,2696,2697,2698,2699,2700,2701],{"id":2507,"depth":195,"text":2508},{"id":2558,"depth":195,"text":2559},{"id":2568,"depth":195,"text":2569},{"id":2575,"depth":195,"text":2576},{"id":2602,"depth":195,"text":2603},{"id":1350,"depth":195,"text":1351},{"id":2659,"depth":195,"text":2660},{"id":2682,"depth":195,"text":2683},"Applebot crawls for search in Siri, Spotlight and Safari, and its data may also train Apple's foundation models. Applebot-Extended is the robots.txt setting that opts out of that training. What each does, how robots.txt applies and how to verify Applebot.",[2704,2707,2710,2713,2716,2719],{"question":2705,"answer":2706},"What is Applebot?","Applebot is Apple's web crawler. Apple says its data powers search features across Apple's ecosystem, including Spotlight, Siri and Safari, and may also be used to help train the Apple foundation models behind generative AI features such as Apple Intelligence.",{"question":2708,"answer":2709},"What is the difference between Applebot and Applebot-Extended?","Applebot crawls web pages. Applebot-Extended doesn't crawl at all: Apple says it is only used to determine how to use the data Applebot crawls. Disallowing Applebot-Extended opts your content out of training Apple's foundation models.",{"question":2711,"answer":2712},"Does blocking Applebot-Extended remove my site from Siri or Spotlight?","No. Apple says pages that disallow Applebot-Extended can still be included in search results, and Applebot keeps crawling them for search in Siri, Spotlight and Safari.",{"question":2714,"answer":2715},"Does Applebot respect robots.txt?","Yes. Apple says Applebot respects standard robots.txt directives that target it, and that if your robots.txt doesn't mention Applebot but mentions Googlebot, Applebot follows the Googlebot rules. It doesn't follow Crawl-delay.",{"question":2717,"answer":2718},"How do I keep my content out of answers generated in Apple products?","Apple says it won't use content tagged nosnippet as additional context when AI models generate output in Apple products and services, and that pages marked isAccessibleForFree false can appear in search results but won't be used that way either. Disallowing Applebot-Extended separately opts your content out of training Apple's foundation models.",{"question":2720,"answer":2721},"How can I tell whether a request really came from Apple?","Run a reverse DNS lookup on the source IP address: genuine Applebot traffic resolves to a host in the applebot.apple.com domain, and a forward lookup of that host should return the same address. Or match the address against the CIDR prefixes in Apple's applebot.json file.",{},7,{"title":2725,"description":2726},"Applebot and Applebot-Extended: What They Do and How to Opt Out","What Applebot crawls for Siri, Spotlight and Safari, how Applebot-Extended opts you out of Apple Intelligence training, how robots.txt applies and how to verify Applebot.",{"loc":859},"ai-crawlers\u002Fapplebot","Applebot powers search in Siri, Spotlight and Safari, and what it crawls may also train Apple's foundation models. Applebot-Extended doesn't crawl: it's the robots.txt setting that opts your content out of that training while you stay in Apple's search.","ukAlMgGBDFETkaHeQmYbc31YmUJDeE_g54shvV6WKXc",{"id":2732,"title":2733,"body":2734,"checked":1949,"description":2921,"documented":204,"extension":205,"faq":2922,"meta":2941,"name":647,"navigation":204,"order":2942,"path":866,"seo":2943,"sitemap":2946,"stem":2947,"tagline":2948,"__hash__":2949},"aiCrawlers\u002Fai-crawlers\u002Fmeta-externalagent.md","Meta-ExternalAgent: Meta's AI crawler and how to control it",{"type":8,"value":2735,"toc":2912},[2736,2739,2743,2818,2822,2825,2828,2834,2838,2841,2845,2848,2854,2861,2869,2871,2874,2880,2884,2887,2893,2900,2904,2907],[11,2737,2738],{},"Meta runs five crawlers, each with its own robots.txt name. Meta says its crawlers may cache your robots.txt for up to 24 hours, so allow that long for a change to take effect.",[15,2740,2742],{"id":2741},"metas-crawlers-at-a-glance","Meta's crawlers at a glance",[20,2744,2745,2756],{},[23,2746,2747],{},[26,2748,2749,2752,2754],{},[29,2750,2751],{},"Crawler",[29,2753,34],{},[29,2755,1638],{},[39,2757,2758,2770,2782,2793,2805],{},[26,2759,2760,2764,2767],{},[44,2761,2762],{},[47,2763,647],{},[44,2765,2766],{},"Crawls the web for use cases such as training foundation AI models or improving products by indexing content directly",[44,2768,2769],{},"Blocked by a Disallow for it",[26,2771,2772,2776,2779],{},[44,2773,2774],{},[47,2775,680],{},[44,2777,2778],{},"Fetches individual links at a user's request, including to help AI complete tasks for users",[44,2780,2781],{},"May bypass robots.txt, because a user requested the fetch",[26,2783,2784,2788,2791],{},[44,2785,2786],{},[47,2787,664],{},[44,2789,2790],{},"Improves Meta AI's search results. Meta says allowing it helps Meta cite and link to your content in Meta AI's responses",[44,2792,2769],{},[26,2794,2795,2800,2803],{},[44,2796,2797],{},[47,2798,2799],{},"Meta-ExternalAds",[44,2801,2802],{},"Crawls for use cases such as improving advertising and other business products",[44,2804,2769],{},[26,2806,2807,2812,2815],{},[44,2808,2809],{},[47,2810,2811],{},"FacebookExternalHit",[44,2813,2814],{},"Fetches pages shared on Facebook, Instagram or Messenger to show their title, description and thumbnail",[44,2816,2817],{},"May bypass robots.txt for security or integrity checks",[15,2819,2821],{"id":2820},"meta-externalagent-and-ai-training","Meta-ExternalAgent and AI training",[11,2823,2824],{},"Meta says Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. It's the Meta crawler to disallow if you don't want your content crawled for those uses.",[11,2826,2827],{},"The rule names only that crawler. It doesn't reach Meta-WebIndexer, which affects whether Meta AI cites and links to you, or FacebookExternalHit, which builds link previews.",[11,2829,2830,2831,2833],{},"Meta's AI crawlers are also busy. Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed, more than Google and OpenAI combined. Meta doesn't document ",[47,2832,102],{}," support or any other way to slow its crawlers down.",[15,2835,2837],{"id":2836},"meta-externalfetcher-fetches-for-users","Meta-ExternalFetcher: fetches for users",[11,2839,2840],{},"Meta-ExternalFetcher fetches individual links at a user's request and supports functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because the user requested the fetch, so keeping it out takes a firewall rule.",[15,2842,2844],{"id":2843},"robotstxt-rules-for-metas-crawlers","robots.txt rules for Meta's crawlers",[11,2846,2847],{},"Keep Meta-ExternalAgent out while Meta AI can still cite you and link previews keep working:",[115,2849,2852],{"className":2850,"code":2851,"language":120},[118],"User-agent: Meta-ExternalAgent\nDisallow: \u002F\n\nUser-agent: Meta-WebIndexer\nAllow: \u002F\n",[47,2853,2851],{"__ignoreMap":123},[11,2855,2856,2857,2860],{},"Meta's user agents are written in lowercase (",[47,2858,2859],{},"meta-externalagent","), but the robots.txt standard (RFC 9309) requires crawlers to match names case-insensitively, so either spelling works.",[11,2862,144,2863,148,2865,152,2867,156],{},[47,2864,147],{},[47,2866,151],{},[47,2868,155],{},[15,2870,1351],{"id":1350},[11,2872,2873],{},"Meta publishes each string with and without the URL part:",[115,2875,2878],{"className":2876,"code":2877,"language":120},[118],"Meta-ExternalAgent:\nmeta-externalagent\u002F1.1 (+https:\u002F\u002Fdevelopers.facebook.com\u002Fdocs\u002Fsharing\u002Fwebmasters\u002Fcrawler)\nmeta-externalagent\u002F1.1\n\nMeta-ExternalFetcher:\nmeta-externalfetcher\u002F1.1 (+https:\u002F\u002Fdevelopers.facebook.com\u002Fdocs\u002Fsharing\u002Fwebmasters\u002Fcrawler)\nmeta-externalfetcher\u002F1.1\n\nMeta-WebIndexer:\nmeta-webindexer\u002F1.1 (+https:\u002F\u002Fdevelopers.facebook.com\u002Fdocs\u002Fsharing\u002Fwebmasters\u002Fcrawler)\nmeta-webindexer\u002F1.1\n\nFacebookExternalHit:\nfacebookexternalhit\u002F1.1 (+http:\u002F\u002Fwww.facebook.com\u002Fexternalhit_uatext.php)\nfacebookexternalhit\u002F1.1\n",[47,2879,2877],{"__ignoreMap":123},[15,2881,2883],{"id":2882},"how-to-verify-a-meta-request","How to verify a Meta request",[11,2885,2886],{},"Meta says a crawler comes from Meta when its source IP address is on the list this command returns, and notes that these addresses change often:",[115,2888,2891],{"className":2889,"code":2890,"language":120},[118],"whois -h whois.radb.net -- '-i origin AS32934' | grep ^route\n",[47,2892,2890],{"__ignoreMap":123},[11,2894,2895,2896,92],{},"Meta publishes no JSON list and doesn't document a reverse DNS check. Questions go to ",[86,2897,2899],{"href":2898},"mailto:webmasters@meta.com","webmasters@meta.com",[15,2901,2903],{"id":2902},"where-to-see-metas-crawlers","Where to see Meta's crawlers",[11,2905,2906],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Meta's crawlers won't appear there. Your server and CDN logs record every request with its user agent and IP address.",[11,2908,2909,2911],{},[86,2910,192],{"href":191}," recognizes Meta-ExternalAgent, Meta-ExternalFetcher and Meta-WebIndexer by their user agents when your server reports page requests to OneLence. It shows which pages each one requested and the visitors Meta AI sends you, next to the same view for OpenAI, Anthropic, Perplexity and Google.",{"title":123,"searchDepth":195,"depth":195,"links":2913},[2914,2915,2916,2917,2918,2919,2920],{"id":2741,"depth":195,"text":2742},{"id":2820,"depth":195,"text":2821},{"id":2836,"depth":195,"text":2837},{"id":2843,"depth":195,"text":2844},{"id":1350,"depth":195,"text":1351},{"id":2882,"depth":195,"text":2883},{"id":2902,"depth":195,"text":2903},"Meta-ExternalAgent crawls the web for uses such as training Meta's foundation AI models. How it differs from Meta-ExternalFetcher, Meta-WebIndexer and FacebookExternalHit, which ones follow robots.txt, and how to block training without breaking link previews.",[2923,2926,2929,2932,2935,2938],{"question":2924,"answer":2925},"What is Meta-ExternalAgent?","Meta-ExternalAgent is one of Meta's web crawlers. Meta says it crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Its user agent starts with meta-externalagent\u002F1.1.",{"question":2927,"answer":2928},"Should I block Meta-ExternalAgent?","Block it if you don't want Meta to crawl your site for uses such as AI model training. A rule for Meta-ExternalAgent doesn't apply to Meta-WebIndexer, which Meta says helps it cite and link to your content in Meta AI's responses, or to FacebookExternalHit, which builds link previews.",{"question":2930,"answer":2931},"Will blocking Meta-ExternalAgent break Facebook link previews?","No. Link previews on Facebook, Instagram and Messenger come from FacebookExternalHit, a separate crawler with its own robots.txt name.",{"question":2933,"answer":2934},"What is Meta-ExternalFetcher?","Meta-ExternalFetcher fetches individual links at a user's request and supports product functions such as helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because a user requested the fetch.",{"question":2936,"answer":2937},"Why is meta-externalagent crawling my site so much?","Meta's AI crawlers are among the busiest: Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed. Meta doesn't document Crawl-delay support or any other way to set a crawl rate, so disallow the paths you don't want crawled, or write to Meta at webmasters@meta.com.",{"question":2939,"answer":2940},"How can I tell whether a request really came from Meta?","Meta says a crawler whose source IP address is on the list returned by the command whois -h whois.radb.net -- '-i origin AS32934' | grep ^route comes from Meta, and notes that these addresses change often. It publishes no JSON list and no reverse DNS check.",{},8,{"title":2944,"description":2945},"Meta-ExternalAgent: Meta's AI Crawler Explained, and How to Block It","What meta-externalagent and meta-externalfetcher do, how they differ from Meta-WebIndexer and FacebookExternalHit, how robots.txt applies and how to block AI training.",{"loc":866},"ai-crawlers\u002Fmeta-externalagent","Meta-ExternalAgent crawls the web for uses such as training foundation AI models. Meta-ExternalFetcher fetches links when users ask and may bypass robots.txt, Meta-WebIndexer helps Meta AI cite and link to you, and FacebookExternalHit builds link previews.","qkaUpMx_kCjEYI9PbkHTRRRDLsYp9xxrT6j45Aa-6lk",{"id":2951,"title":2952,"body":2953,"checked":1949,"description":3089,"documented":3090,"extension":205,"faq":3091,"meta":3107,"name":698,"navigation":204,"order":3108,"path":872,"seo":3109,"sitemap":3112,"stem":3113,"tagline":3114,"__hash__":3115},"aiCrawlers\u002Fai-crawlers\u002Fbytespider.md","Bytespider: ByteDance's crawler, and why robots.txt may not stop it",{"type":8,"value":2954,"toc":3081},[2955,2958,2962,2991,2995,2998,3018,3021,3025,3028,3034,3040,3043,3045,3048,3054,3060,3064,3067,3071,3076],[11,2956,2957],{},"Bytespider is one of the crawlers site owners ask about most, and one of the least documented. ByteDance publishes no page we could reach that describes it, so everything below comes from independent reports, each named with its date.",[15,2959,2961],{"id":2960},"what-is-known-about-bytespider","What is known about Bytespider",[822,2963,2964,2970,2976,2985],{},[825,2965,2966,2969],{},[913,2967,2968],{},"Operator."," The user agent strings that others record carry ByteDance and Toutiao addresses. ByteDance is TikTok's parent company.",[825,2971,2972,2975],{},[913,2973,2974],{},"Purpose."," Not documented by ByteDance on any page we could reach. Third-party pages claim search and AI training uses without citing ByteDance.",[825,2977,2978,2981,2982,2984],{},[913,2979,2980],{},"robots.txt name."," Everyone uses ",[47,2983,698],{},". Whether the crawler honors it is disputed, as below.",[825,2986,2987,2990],{},[913,2988,2989],{},"IP list."," None from ByteDance that we could reach.",[15,2992,2994],{"id":2993},"does-bytespider-follow-robotstxt","Does Bytespider follow robots.txt?",[11,2996,2997],{},"Several independent reports say it doesn't, at least not reliably:",[822,2999,3000,3006,3012],{},[825,3001,3002,3005],{},[913,3003,3004],{},"Fortune, October 2024."," Fortune reported research showing that Bytespider didn't respect robots.txt, and quoted Kasada's CEO saying it had been scraping data at about 25 times the rate of GPTBot. The same article made a similar claim about OpenAI's and Anthropic's crawlers, whose own documentation says GPTBot and ClaudeBot follow robots.txt.",[825,3007,3008,3011],{},[913,3009,3010],{},"HAProxy, October 2024."," HAProxy wrote that some AI crawlers, Bytespider included, don't identify themselves transparently, try to pretend to be real users and ignore robots.txt. Close to 90% of the AI crawler traffic HAProxy saw came from Bytespider.",[825,3013,3014,3017],{},[913,3015,3016],{},"TollBit, first half of 2026",", as reported by Search Engine Journal in August 2026: ChatGPT-User, Bytespider and YouBot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them.",[11,3019,3020],{},"Its volume has dropped since 2024. Cloudflare reported in July 2024 that Bytespider led the AI bots it saw in requests, in how much of the web it crawled and in how often it was blocked. In July 2025, Cloudflare reported that Bytespider's request volume had fallen 85%, from second to eighth place in crawler share, at 2.9%.",[15,3022,3024],{"id":3023},"how-to-block-bytespider","How to block Bytespider",[11,3026,3027],{},"Start with robots.txt, which costs nothing:",[115,3029,3032],{"className":3030,"code":3031,"language":120},[118],"User-agent: Bytespider\nDisallow: \u002F\n",[47,3033,3031],{"__ignoreMap":123},[11,3035,3036,3037,3039],{},"Because the reports above say robots.txt isn't reliably honored, back it with a firewall or WAF rule that blocks requests whose user agent contains ",[47,3038,698],{},". The rule doesn't depend on the crawler's cooperation, and it also stops impostors that use the name, which does no harm here.",[11,3041,3042],{},"Blocking Bytespider doesn't affect any ByteDance service we know of, since ByteDance documents none that depends on it.",[15,3044,1351],{"id":1350},[11,3046,3047],{},"ByteDance publishes no strings that we could reach. Third parties record different ones, for example:",[115,3049,3052],{"className":3050,"code":3051,"language":120},[118],"Mozilla\u002F5.0 (compatible; Bytespider; spider-feedback@bytedance.com)\n\nMozilla\u002F5.0 (compatible; Bytespider; https:\u002F\u002Fzhanzhang.toutiao.com\u002F) AppleWebKit\u002F537.36 (KHTML, like Gecko) Chrome\u002F70.0.0.0 Safari\u002F537.36\n",[47,3053,3051],{"__ignoreMap":123},[11,3055,3056,3057,3059],{},"Match on ",[47,3058,698],{}," rather than on the whole string.",[15,3061,3063],{"id":3062},"can-you-verify-bytespider","Can you verify Bytespider?",[11,3065,3066],{},"No. Without an IP list or a reverse DNS method from ByteDance, a request that says Bytespider can't be confirmed as genuine. The third-party IP lists that circulate are old or unverified.",[15,3068,3070],{"id":3069},"where-to-see-bytespider","Where to see Bytespider",[11,3072,3073,3074,92],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Bytespider won't appear there. Search your server and CDN logs for ",[47,3075,698],{},[11,3077,3078,3080],{},[86,3079,192],{"href":191}," recognizes Bytespider by its user agent when your server reports page requests to OneLence, and shows which pages it requested and how often, marked as matched by user agent because ByteDance publishes no IP list.",{"title":123,"searchDepth":195,"depth":195,"links":3082},[3083,3084,3085,3086,3087,3088],{"id":2960,"depth":195,"text":2961},{"id":2993,"depth":195,"text":2994},{"id":3023,"depth":195,"text":3024},{"id":1350,"depth":195,"text":1351},{"id":3062,"depth":195,"text":3063},{"id":3069,"depth":195,"text":3070},"Bytespider is the crawler associated with ByteDance, TikTok's parent company. ByteDance publishes no documentation we could reach, so here is what independent reports say about its robots.txt behavior and volume, and how to block it reliably.",false,[3092,3095,3098,3101,3104],{"question":3093,"answer":3094},"What is Bytespider?","Bytespider is a web crawler associated with ByteDance, TikTok's parent company: the user agent strings people record carry ByteDance and Toutiao addresses. ByteDance publishes no documentation we could reach on what Bytespider collects or why.",{"question":3096,"answer":3097},"Does Bytespider respect robots.txt?","Not reliably, according to several independent reports. Fortune and HAProxy reported in 2024 that it ignored robots.txt, and TollBit's report on the first half of 2026 found Bytespider accessing disallowed pages on nearly half of the European sites that had explicitly listed it. ByteDance publishes no statement we could check.",{"question":3099,"answer":3100},"How do I block Bytespider?","Add a robots.txt group with User-agent: Bytespider and Disallow: \u002F, and back it with a firewall or WAF rule that blocks requests whose user agent contains Bytespider. The firewall rule doesn't depend on the crawler's cooperation, and it also stops impostors that use the name.",{"question":3102,"answer":3103},"Is Bytespider used to train AI?","ByteDance doesn't say on any page we could reach. Third-party pages claim search and AI training uses, but they cite no ByteDance statement, so treat those claims as unconfirmed.",{"question":3105,"answer":3106},"Can I verify that a request really came from Bytespider?","No. ByteDance publishes no IP list and no reverse DNS method that we could reach, and the third-party IP lists that circulate are old or unverified. A request that says Bytespider could come from anyone.",{},9,{"title":3110,"description":3111},"Bytespider: ByteDance's Crawler, and How to Block It","What is known about Bytespider, ByteDance's crawler: its user agent, reports that it ignores robots.txt, its crawl volume, and how to block it with robots.txt and a firewall.",{"loc":872},"ai-crawlers\u002Fbytespider","Bytespider is the crawler associated with ByteDance, TikTok's parent company. ByteDance publishes no documentation we could reach, and several independent reports say it doesn't reliably follow robots.txt, so a firewall rule is the dependable way to block it.","9Zgj7RhVHVHLd5iL8XfTPPcBu17V4hH4HdZqmDBbRzM",{"id":3117,"title":3118,"body":3119,"checked":1949,"description":3329,"documented":204,"extension":205,"faq":3330,"meta":3349,"name":717,"navigation":204,"order":3350,"path":878,"seo":3351,"sitemap":3354,"stem":3355,"tagline":3356,"__hash__":3357},"aiCrawlers\u002Fai-crawlers\u002Fccbot.md","CCBot: Common Crawl's crawler, and what blocking it changes",{"type":8,"value":3120,"toc":3321},[3121,3124,3128,3198,3202,3205,3228,3232,3235,3241,3244,3250,3258,3262,3282,3285,3289,3296,3302,3309,3313,3316],[11,3122,3123],{},"Common Crawl is a nonprofit that crawls the web and gives its archive away. It publishes a crawl about once a month, each typically with more than two billion web pages, and says the archive has become one of the most widely used sources of training data for large language models. CCBot is the crawler that collects it.",[15,3125,3127],{"id":3126},"ccbot-at-a-glance","CCBot at a glance",[20,3129,3130,3139],{},[23,3131,3132],{},[26,3133,3134,3137],{},[29,3135,3136],{},"Detail",[29,3138,717],{},[39,3140,3141,3149,3159,3168,3177,3186],{},[26,3142,3143,3146],{},[44,3144,3145],{},"Operator",[44,3147,3148],{},"Common Crawl Foundation, a 501(c)(3) nonprofit",[26,3150,3151,3154],{},[44,3152,3153],{},"User agent",[44,3155,3156],{},[47,3157,3158],{},"CCBot\u002F2.0 (https:\u002F\u002Fcommoncrawl.org\u002Ffaq\u002F)",[26,3160,3161,3164],{},[44,3162,3163],{},"robots.txt name",[44,3165,3166],{},[47,3167,717],{},[26,3169,3170,3172],{},[44,3171,326],{},[44,3173,3174,3175],{},"Yes, including ",[47,3176,102],{},[26,3178,3179,3181],{},[44,3180,329],{},[44,3182,3183],{},[86,3184,730],{"href":728,"rel":3185},[90],[26,3187,3188,3191],{},[44,3189,3190],{},"Reverse DNS",[44,3192,3193,3194,3197],{},"Hosts under ",[47,3195,3196],{},"crawl.commoncrawl.org",", except over IPv6",[15,3199,3201],{"id":3200},"how-ccbot-treats-robotstxt","How CCBot treats robots.txt",[11,3203,3204],{},"Common Crawl says:",[822,3206,3207,3210,3216,3222,3225],{},[825,3208,3209],{},"A CCBot disallow in your robots.txt stops its crawler crawling your site.",[825,3211,3212,3213,3215],{},"It obeys ",[47,3214,102],{},": a larger number tells CCBot to slow down.",[825,3217,3218,3219,3221],{},"It honors ",[47,3220,90],{}," on the links on your site.",[825,3223,3224],{},"It periodically checks whether your robots.txt has changed, without saying how often.",[825,3226,3227],{},"It slows down on its own when your server answers with HTTP 429 or 5xx errors.",[15,3229,3231],{"id":3230},"robotstxt-rules-for-ccbot","robots.txt rules for CCBot",[11,3233,3234],{},"Block it:",[115,3236,3239],{"className":3237,"code":3238,"language":120},[118],"User-agent: CCBot\nDisallow: \u002F\n",[47,3240,3238],{"__ignoreMap":123},[11,3242,3243],{},"Slow it down instead, as in Common Crawl's example:",[115,3245,3248],{"className":3246,"code":3247,"language":120},[118],"User-agent: CCBot\nCrawl-delay: 2\n",[47,3249,3247],{"__ignoreMap":123},[11,3251,144,3252,148,3254,152,3256,156],{},[47,3253,147],{},[47,3255,151],{},[47,3257,155],{},[15,3259,3261],{"id":3260},"what-blocking-ccbot-changes-and-what-it-doesnt","What blocking CCBot changes, and what it doesn't",[822,3263,3264,3270,3276],{},[825,3265,3266,3269],{},[913,3267,3268],{},"It stops future crawls."," That's the effect Common Crawl documents.",[825,3271,3272,3275],{},[913,3273,3274],{},"It doesn't remove pages already archived."," Common Crawl doesn't document removing pages from crawls it has published, and copies that others downloaded aren't affected either.",[825,3277,3278,3281],{},[913,3279,3280],{},"Legal requests are a separate route."," Common Crawl says it processes legal requests as it receives them, and publishes the opt-out requests in an Opt-out Registry to alert the people who use its data.",[11,3283,3284],{},"Common Crawl argues the other side on its own blog: \"If you're blocked at the crawl layer, you're excluded entirely.\" That's its view, and the choice is yours.",[15,3286,3288],{"id":3287},"how-to-verify-ccbot","How to verify CCBot",[11,3290,3291,3292,3295],{},"Common Crawl runs CCBot on dedicated IP ranges with reverse DNS, except over IPv6, where reverse DNS isn't supported yet. A genuine request's IP address resolves to a host such as ",[47,3293,3294],{},"18-97-14-84.crawl.commoncrawl.org",". Common Crawl's own examples:",[115,3297,3300],{"className":3298,"code":3299,"language":120},[118],"host 18.97.14.84\ndig -x 18.97.14.84\ndig 18-97-14-84.crawl.commoncrawl.org A\n",[47,3301,3299],{"__ignoreMap":123},[11,3303,3304,3305,3308],{},"You can also check the address against ",[86,3306,730],{"href":728,"rel":3307},[90],". Common Crawl warns that other crawlers falsely identify themselves as CCBot, so don't trust the user agent alone.",[15,3310,3312],{"id":3311},"where-to-see-ccbot","Where to see CCBot",[11,3314,3315],{},"Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so CCBot won't appear there. Your server and CDN logs record every request with its user agent and IP address.",[11,3317,3318,3320],{},[86,3319,192],{"href":191}," recognizes CCBot when your server reports page requests to OneLence, marks the requests confirmed by Common Crawl's IP list as verified, and shows which pages it read and how often, next to the crawlers of OpenAI, Anthropic, Perplexity and Google.",{"title":123,"searchDepth":195,"depth":195,"links":3322},[3323,3324,3325,3326,3327,3328],{"id":3126,"depth":195,"text":3127},{"id":3200,"depth":195,"text":3201},{"id":3230,"depth":195,"text":3231},{"id":3260,"depth":195,"text":3261},{"id":3287,"depth":195,"text":3288},{"id":3311,"depth":195,"text":3312},"CCBot is the crawler of Common Crawl, the nonprofit whose free web archive is one of the most widely used sources of training data for large language models. What it does, how it follows robots.txt and Crawl-delay, how to verify it, and what blocking it changes.",[3331,3334,3337,3340,3343,3346],{"question":3332,"answer":3333},"What is CCBot?","CCBot is the web crawler of Common Crawl, a nonprofit that crawls the web and freely provides its archives and datasets to the public. Its user agent is CCBot\u002F2.0 (https:\u002F\u002Fcommoncrawl.org\u002Ffaq\u002F).",{"question":3335,"answer":3336},"Is Common Crawl used to train AI models?","Yes. Common Crawl says its archive has been cited in over 12,000 research papers and has become one of the most widely used sources of training data for large language models. The archive is free to download, so anyone can use it.",{"question":3338,"answer":3339},"Does CCBot respect robots.txt?","Yes. Common Crawl says adding User-agent: CCBot and Disallow: \u002F to your robots.txt stops its crawler, that it obeys Crawl-delay and that it honors nofollow on links. It also slows down when your server answers with HTTP 429 or 5xx errors.",{"question":3341,"answer":3342},"Does blocking CCBot remove my pages from Common Crawl?","Blocking stops future crawls. Common Crawl doesn't document removing pages from crawls it has already published, and copies that others have downloaded aren't affected. It processes legal requests and publishes the opt-out requests it receives in an Opt-out Registry.",{"question":3344,"answer":3345},"How do I slow CCBot down?","Add Crawl-delay to a CCBot group in your robots.txt, as in Common Crawl's own example: User-agent: CCBot, then Crawl-delay: 2. A larger number tells CCBot to crawl more slowly.",{"question":3347,"answer":3348},"How can I tell whether a request really came from Common Crawl?","Run a reverse DNS lookup on the source IP address: CCBot runs on dedicated IP ranges whose hosts are under crawl.commoncrawl.org, except over IPv6. Or check the address against index.commoncrawl.org\u002Fccbot.json. Common Crawl warns that other crawlers falsely identify themselves as CCBot.",{},10,{"title":3352,"description":3353},"CCBot: Common Crawl's Crawler, AI Training, and How to Block It","What CCBot is, why Common Crawl matters for AI training, how it follows robots.txt and Crawl-delay, how to verify it, and what blocking it changes for pages already archived.",{"loc":878},"ai-crawlers\u002Fccbot","CCBot builds Common Crawl's free web archive, which Common Crawl calls one of the most widely used sources of training data for large language models. It follows robots.txt and Crawl-delay, but blocking it only stops future crawls.","MEzpD7PPqYOrNYi2xQw771PUioTCr_40c2cMlvAtLpM",1791487629007]