Amazon documents three agents on one page and says each user agent setting is independent of the others. Amazon says changes can take about 24 hours to reach its systems, though the same page also says the crawlers may use a copy of your robots.txt cached within the last 30 days.
Amazon's crawlers at a glance
| Agent | What it does | What robots.txt does | IP list |
|---|---|---|---|
Amazonbot | Improves Amazon's products and services, and may be used to train Amazon AI models | A Disallow stops Amazonbot crawling those pages | Amazonbot IP addresses |
Amzn-SearchBot | Makes content eligible for search experiences such as Alexa. Not used for generative AI training | A Disallow stops it crawling. Amazon says allowing it is what makes your content eligible | Amzn-SearchBot IP addresses |
Amzn-User | Supports user actions, such as answering Alexa questions that need up-to-date information. Not used for generative AI training | May not follow every directive, because a user can start its actions | Amzn-User IP addresses |
How Amazon's crawlers treat robots.txt
Amazon says its crawlers respect the Robots Exclusion Protocol, honoring the user-agent line and the allow and disallow directives. In detail:
- They fetch each host's robots.txt, or use a copy cached within the last 30 days. When the file can't be fetched, they behave as if it doesn't exist.
- They honor the rules each host of a domain exposes, so every subdomain needs its own robots.txt.
- They don't support
Crawl-delay, and Amazon documents no other way to slow them down. - They respect
rel=nofollowon links and the page-level robots meta tagsnoarchive(do not use the page for model training),noindexandnone. - If robots.txt doesn't mention Amzn-SearchBot but allows other search bots, Amzn-SearchBot follows the rules given to those search bots.
robots.txt rules for Amazon's crawlers
Opt out of Amazonbot, AI training included, while staying eligible for Amazon's search experiences such as Alexa:
User-agent: Amazonbot
Disallow: /
User-agent: Amzn-SearchBot
Allow: /
Opt out of all three:
User-agent: Amazonbot
Disallow: /
User-agent: Amzn-SearchBot
Disallow: /
User-agent: Amzn-User
Disallow: /
Amzn-User may still fetch a page when a user starts the action. To keep a single page out of training without blocking any crawler, Amazon honors the noarchive robots meta tag, which it defines as "do not use the page for model training":
<meta name="robots" content="noarchive">
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
User agent strings
Amazon publishes these strings. Chrome/W.X.Y.Z stands for a Chrome version, so match on the agent's name rather than the whole string:
Amazonbot:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36
Amzn-SearchBot:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36
Amzn-User:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36
How to verify an Amazon request
Any script can send these strings. Amazon publishes the IP addresses each agent uses, linked in the table above, as lists of single addresses. Check a request's source IP address against the list for its agent. Amazon doesn't document a reverse DNS check, so don't rely on older guides that describe one. Publishers with questions can write to [email protected].
Where to see Amazon's crawlers
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Amazonbot won't appear there. Your server and CDN logs record every request with its user agent and IP address. Search them for Amazonbot, Amzn-SearchBot and Amzn-User.
OneLence AI visibility does this for you when your server reports page requests to OneLence. It recognizes Amazonbot, Amzn-SearchBot and Amzn-User by their user agents and shows which pages each one requested, next to the same view for OpenAI, Anthropic, Perplexity and Google.
Frequently asked questions
What is Amazonbot?
Amazonbot is Amazon's web crawler. Amazon says it is used to improve Amazon's products and services, helps provide more accurate information to customers, and may be used to train Amazon AI models.
Does Amazonbot respect robots.txt?
Yes. Amazon says Amazonbot, Amzn-SearchBot and Amzn-User respect the Robots Exclusion Protocol, honoring the user-agent line and the allow and disallow directives. They fetch each host's robots.txt or use a copy cached within the last 30 days. They don't support Crawl-delay, and Amzn-User may not follow every directive because a user can start its actions.
Should I block Amazonbot?
Block it if you don't want your content used to train Amazon AI models. A rule for Amazonbot doesn't apply to Amzn-SearchBot, which makes your content eligible for search experiences such as Alexa, or to Amzn-User, which fetches pages for live requests. Amazon says each user agent setting is independent of the others.
How do I slow Amazonbot down?
robots.txt can't do it: Amazon says its crawlers don't support the Crawl-delay directive, and it documents no other way to set a crawl rate. You can disallow the paths you don't want crawled, or write to Amazon at [email protected].
How can I tell whether a request really came from Amazon?
Check the source IP address against the address list Amazon publishes for that agent on its developer site: one each for Amazonbot, Amzn-SearchBot and Amzn-User. Amazon doesn't document a reverse DNS check, so don't rely on older guides that describe one.
