Perplexity runs two agents, and each has its own robots.txt setting: PerplexityBot builds the search index, and Perplexity-User fetches pages for people's questions. Perplexity says each setting works independently and that its systems can take up to 24 hours to reflect a change.
Perplexity's crawlers at a glance
| Agent | What it does | What robots.txt does | IP list |
|---|---|---|---|
PerplexityBot | Surfaces and links websites in Perplexity's search results. Not used to crawl content for AI foundation models | A Disallow stops indexing of the page's full or partial text. The domain, headline and a brief factual summary may still be indexed | perplexitybot.json |
Perplexity-User | Visits a page when someone's question in Perplexity needs it. Not used for web crawling or training | Generally ignored, because a user requested the fetch | perplexity-user.json |
Perplexity also says it works with third-party crawlers to help build its search index, and that its agreements require them to respect robots.txt. It doesn't name them or their user agents, so you can't write rules for them.
PerplexityBot: Perplexity's search index
PerplexityBot finds the pages that Perplexity's search results show and link to. Perplexity says it only crawls content in compliance with robots.txt, and it recommends allowing PerplexityBot, and requests from its published IP ranges, if you want your site to appear in Perplexity's search results.
Blocking it doesn't make a page disappear entirely. Perplexity says it won't index the full or partial text of a page that disallows PerplexityBot, but it may still index the domain, the headline and a brief factual summary.
Perplexity-User: the visit behind an answer
When someone asks Perplexity a question, it might visit a web page with the Perplexity-User agent to help give an accurate answer. Perplexity says that since a user requested the fetch, Perplexity-User generally ignores robots.txt rules. Keeping it out takes a firewall rule, which also keeps your pages out of the answers people ask for.
A Perplexity-User request in your logs means someone's question needed that page at that moment. It shows that Perplexity read the page, not whether the answer quoted or linked it.
Does Perplexity train on your content?
Perplexity says no. It doesn't build foundation models, so your content won't be used for AI model pre-training. There is no Perplexity training crawler, which is why the training opt-out rules for OpenAI, Anthropic and Google have no Perplexity line.
robots.txt rules for Perplexity
Stay in Perplexity's search results. This is also what happens when your robots.txt doesn't mention PerplexityBot:
User-agent: PerplexityBot
Allow: /
Keep your pages' text out of Perplexity's index:
User-agent: PerplexityBot
Disallow: /
A rule for Perplexity-User is generally ignored, so robots.txt can't keep it out. Perplexity doesn't document support for Crawl-delay.
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
The 2025 dispute with Cloudflare
In August 2025, Cloudflare reported that when PerplexityBot was blocked, Perplexity also crawled with an undeclared user agent that impersonated Chrome on macOS, from IP addresses not on Perplexity's lists. Cloudflare removed Perplexity from its verified bots. Perplexity replied the same day that Cloudflare had misattributed the traffic, which it said came from BrowserBase's automated browser service. The two accounts still disagree.
User agent strings
Perplexity publishes these strings:
PerplexityBot:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-User:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
How to verify a Perplexity request
Any script can send these strings. To confirm a request came from Perplexity, check its source IP address against the list for that agent in the table above. Perplexity recommends combining user agent matching with IP address verification when you write firewall rules for its bots. It doesn't document a reverse DNS check.
Where to see Perplexity's crawlers and the visitors it sends
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so PerplexityBot and Perplexity-User won't appear there. Your server and CDN logs record every request with its user agent and IP address. People who click a link in a Perplexity answer are different: they are visitors, and when their browser passes the referrer they show up in your analytics with perplexity.ai as the source.
OneLence AI visibility puts both sides in one report when your server reports page requests to OneLence. It marks PerplexityBot and Perplexity-User requests confirmed by Perplexity's IP lists as verified, and shows which pages PerplexityBot read, which pages Perplexity-User fetched while answering people, and the visitors Perplexity sent you and the ones that converted.
Frequently asked questions
What is PerplexityBot?
PerplexityBot is Perplexity's search crawler. Perplexity says it is designed to surface and link websites in search results on Perplexity, and that it isn't used to crawl content for AI foundation models.
What is the difference between PerplexityBot and Perplexity-User?
PerplexityBot crawls the web for Perplexity's search index and follows robots.txt. Perplexity-User visits a page when someone asks Perplexity a question that needs it. Perplexity says that because a user requested the fetch, Perplexity-User generally ignores robots.txt rules.
Does Perplexity respect robots.txt?
PerplexityBot does: Perplexity says it only crawls content in compliance with robots.txt. Perplexity-User generally doesn't, because a person asked for the page. In 2025 Cloudflare reported undeclared crawling that it attributed to Perplexity, and Perplexity denied it.
Will blocking PerplexityBot remove my site from Perplexity?
Not completely. Perplexity says it won't index the full or partial text of a page that disallows PerplexityBot, but it may still index the domain, the headline and a brief factual summary. Perplexity-User can also still visit when someone asks about your page, since it generally ignores robots.txt.
Does Perplexity train AI models on my content?
Perplexity says no. It doesn't build foundation models, so your content won't be used for AI model pre-training. That's why there is no Perplexity training crawler to block.
How can I tell whether a request really came from Perplexity?
Check the source IP address against the list Perplexity publishes for that agent: perplexity.com/perplexitybot.json for PerplexityBot and perplexity.com/perplexity-user.json for Perplexity-User. Perplexity recommends combining user agent matching with IP address checks in firewall rules.
