Meta runs five crawlers, each with its own robots.txt name. Meta says its crawlers may cache your robots.txt for up to 24 hours, so allow that long for a change to take effect.
Meta's crawlers at a glance
| Crawler | What it does | robots.txt |
|---|---|---|
Meta-ExternalAgent | Crawls the web for use cases such as training foundation AI models or improving products by indexing content directly | Blocked by a Disallow for it |
Meta-ExternalFetcher | Fetches individual links at a user's request, including to help AI complete tasks for users | May bypass robots.txt, because a user requested the fetch |
Meta-WebIndexer | Improves Meta AI's search results. Meta says allowing it helps Meta cite and link to your content in Meta AI's responses | Blocked by a Disallow for it |
Meta-ExternalAds | Crawls for use cases such as improving advertising and other business products | Blocked by a Disallow for it |
FacebookExternalHit | Fetches pages shared on Facebook, Instagram or Messenger to show their title, description and thumbnail | May bypass robots.txt for security or integrity checks |
Meta-ExternalAgent and AI training
Meta says Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. It's the Meta crawler to disallow if you don't want your content crawled for those uses.
The rule names only that crawler. It doesn't reach Meta-WebIndexer, which affects whether Meta AI cites and links to you, or FacebookExternalHit, which builds link previews.
Meta's AI crawlers are also busy. Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed, more than Google and OpenAI combined. Meta doesn't document Crawl-delay support or any other way to slow its crawlers down.
Meta-ExternalFetcher: fetches for users
Meta-ExternalFetcher fetches individual links at a user's request and supports functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because the user requested the fetch, so keeping it out takes a firewall rule.
robots.txt rules for Meta's crawlers
Keep Meta-ExternalAgent out while Meta AI can still cite you and link previews keep working:
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Meta-WebIndexer
Allow: /
Meta's user agents are written in lowercase (meta-externalagent), but the robots.txt standard (RFC 9309) requires crawlers to match names case-insensitively, so either spelling works.
A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.
User agent strings
Meta publishes each string with and without the URL part:
Meta-ExternalAgent:
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-externalagent/1.1
Meta-ExternalFetcher:
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-externalfetcher/1.1
Meta-WebIndexer:
meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-webindexer/1.1
FacebookExternalHit:
facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)
facebookexternalhit/1.1
How to verify a Meta request
Meta says a crawler comes from Meta when its source IP address is on the list this command returns, and notes that these addresses change often:
whois -h whois.radb.net -- '-i origin AS32934' | grep ^route
Meta publishes no JSON list and doesn't document a reverse DNS check. Questions go to [email protected].
Where to see Meta's crawlers
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Meta's crawlers won't appear there. Your server and CDN logs record every request with its user agent and IP address.
OneLence AI visibility recognizes Meta-ExternalAgent, Meta-ExternalFetcher and Meta-WebIndexer by their user agents when your server reports page requests to OneLence. It shows which pages each one requested and the visitors Meta AI sends you, next to the same view for OpenAI, Anthropic, Perplexity and Google.
Frequently asked questions
What is Meta-ExternalAgent?
Meta-ExternalAgent is one of Meta's web crawlers. Meta says it crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Its user agent starts with meta-externalagent/1.1.
Should I block Meta-ExternalAgent?
Block it if you don't want Meta to crawl your site for uses such as AI model training. A rule for Meta-ExternalAgent doesn't apply to Meta-WebIndexer, which Meta says helps it cite and link to your content in Meta AI's responses, or to FacebookExternalHit, which builds link previews.
Will blocking Meta-ExternalAgent break Facebook link previews?
No. Link previews on Facebook, Instagram and Messenger come from FacebookExternalHit, a separate crawler with its own robots.txt name.
What is Meta-ExternalFetcher?
Meta-ExternalFetcher fetches individual links at a user's request and supports product functions such as helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because a user requested the fetch.
Why is meta-externalagent crawling my site so much?
Meta's AI crawlers are among the busiest: Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed. Meta doesn't document Crawl-delay support or any other way to set a crawl rate, so disallow the paths you don't want crawled, or write to Meta at [email protected].
How can I tell whether a request really came from Meta?
Meta says a crawler whose source IP address is on the list returned by the command whois -h whois.radb.net -- '-i origin AS32934' | grep ^route comes from Meta, and notes that these addresses change often. It publishes no JSON list and no reverse DNS check.
