Apple runs one crawler, Applebot, and one extra robots.txt setting, Applebot-Extended, that controls how Applebot's data may be used for AI.
Applebot and Applebot-Extended at a glance
| Name | What it is | What a Disallow does |
|---|---|---|
Applebot | Apple's crawler. Its data powers search in Spotlight, Siri and Safari, and may help train Apple foundation models | Stops Applebot crawling those pages, for search and training alike |
Applebot-Extended | A robots.txt setting, not a crawler. It decides whether Applebot's data may train Apple's foundation models | Opts those pages out of training. They can still appear in search results |
Apple publishes Applebot's IP ranges in applebot.json.
What Applebot does
Apple says the data Applebot crawls powers features such as the search technology in Spotlight, Siri and Safari. The same data may also help train the Apple foundation models behind generative AI features across Apple products, including Apple Intelligence, Services and Developer Tools.
Applebot may render pages in a browser. Apple says that if robots.txt blocks your JavaScript, CSS or other resources, Applebot may not be able to render the content properly.
Applebot-Extended: the training opt-out
Applebot-Extended doesn't crawl webpages. Apple says it is only used to determine how to use the data Applebot crawls, and that disallowing it opts your content out of being used to train Apple's general purpose foundation models. Pages that disallow Applebot-Extended can still be included in search results.
How Applebot treats robots.txt
- Applebot respects standard robots.txt directives that target it.
- If your robots.txt doesn't mention Applebot but mentions Googlebot, Applebot follows the Googlebot rules. A site that blocks Googlebot and never names Applebot blocks Applebot too.
- Applebot doesn't follow
Crawl-delay. - It supports robots meta tags in HTML and indexing directives in the
X-Robots-TagHTTP header.
robots.txt rules for Apple
Stay in search in Siri, Spotlight and Safari, but opt out of training Apple's foundation models:
User-agent: Applebot-Extended
Disallow: /
Opt out of training for one folder only, as in Apple's own example:
User-agent: Applebot-Extended
Disallow: /private/
To keep a page in search but out of the context AI models use when they generate output in Apple products, Apple names two signals: it won't use content tagged nosnippet that way, nor pages marked isAccessibleForFree: false.
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
User agent strings
Apple gives these as examples for desktop and mobile. The stable part is Applebot/0.1; +http://www.apple.com/go/applebot, so match on that:
Desktop (example):
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)
Mobile (example):
Mozilla/5.0 (iPhone; CPU iPhone OS 17_4_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4.1 Mobile/15E148 Safari/604.1 (Applebot/0.1; +http://www.apple.com/go/applebot)
Applebot-Extended has no user agent string, because it never sends a request.
How to verify Applebot
Apple says Applebot traffic is generally identified by reverse DNS in the applebot.apple.com domain. Apple's example:
host 17.58.101.179
# points to 17-58-101-179.applebot.apple.com
host 17-58-101-179.applebot.apple.com
# should return 17.58.101.179
You can also match the IP address against the CIDR prefixes in applebot.json.
Where to see Applebot
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Applebot won't appear there. Your server and CDN logs record its requests.
OneLence AI visibility counts Applebot as a search crawler, like Googlebot and Bingbot, rather than as AI, and confirms its requests against Apple's IP list. Because Applebot can render pages, the OneLence tracking tag can record it even without server-side tracking.
Frequently asked questions
What is Applebot?
Applebot is Apple's web crawler. Apple says its data powers search features across Apple's ecosystem, including Spotlight, Siri and Safari, and may also be used to help train the Apple foundation models behind generative AI features such as Apple Intelligence.
What is the difference between Applebot and Applebot-Extended?
Applebot crawls web pages. Applebot-Extended doesn't crawl at all: Apple says it is only used to determine how to use the data Applebot crawls. Disallowing Applebot-Extended opts your content out of training Apple's foundation models.
Does blocking Applebot-Extended remove my site from Siri or Spotlight?
No. Apple says pages that disallow Applebot-Extended can still be included in search results, and Applebot keeps crawling them for search in Siri, Spotlight and Safari.
Does Applebot respect robots.txt?
Yes. Apple says Applebot respects standard robots.txt directives that target it, and that if your robots.txt doesn't mention Applebot but mentions Googlebot, Applebot follows the Googlebot rules. It doesn't follow Crawl-delay.
How do I keep my content out of answers generated in Apple products?
Apple says it won't use content tagged nosnippet as additional context when AI models generate output in Apple products and services, and that pages marked isAccessibleForFree false can appear in search results but won't be used that way either. Disallowing Applebot-Extended separately opts your content out of training Apple's foundation models.
How can I tell whether a request really came from Apple?
Run a reverse DNS lookup on the source IP address: genuine Applebot traffic resolves to a host in the applebot.apple.com domain, and a forward lookup of that host should return the same address. Or match the address against the CIDR prefixes in Apple's applebot.json file.
