AI crawler guide

Meta-ExternalAgent: Meta's AI crawler and how to control it

Meta-ExternalAgent crawls the web for uses such as training foundation AI models. Meta-ExternalFetcher fetches links when users ask and may bypass robots.txt, Meta-WebIndexer helps Meta AI cite and link to you, and FacebookExternalHit builds link previews.

Last checked against official documentation on October 5, 2026.

Meta runs five crawlers, each with its own robots.txt name. Meta says its crawlers may cache your robots.txt for up to 24 hours, so allow that long for a change to take effect.

Meta's crawlers at a glance

CrawlerWhat it doesrobots.txt
Meta-ExternalAgentCrawls the web for use cases such as training foundation AI models or improving products by indexing content directlyBlocked by a Disallow for it
Meta-ExternalFetcherFetches individual links at a user's request, including to help AI complete tasks for usersMay bypass robots.txt, because a user requested the fetch
Meta-WebIndexerImproves Meta AI's search results. Meta says allowing it helps Meta cite and link to your content in Meta AI's responsesBlocked by a Disallow for it
Meta-ExternalAdsCrawls for use cases such as improving advertising and other business productsBlocked by a Disallow for it
FacebookExternalHitFetches pages shared on Facebook, Instagram or Messenger to show their title, description and thumbnailMay bypass robots.txt for security or integrity checks

Meta-ExternalAgent and AI training

Meta says Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. It's the Meta crawler to disallow if you don't want your content crawled for those uses.

The rule names only that crawler. It doesn't reach Meta-WebIndexer, which affects whether Meta AI cites and links to you, or FacebookExternalHit, which builds link previews.

Meta's AI crawlers are also busy. Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed, more than Google and OpenAI combined. Meta doesn't document Crawl-delay support or any other way to slow its crawlers down.

Meta-ExternalFetcher: fetches for users

Meta-ExternalFetcher fetches individual links at a user's request and supports functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because the user requested the fetch, so keeping it out takes a firewall rule.

robots.txt rules for Meta's crawlers

Keep Meta-ExternalAgent out while Meta AI can still cite you and link previews keep working:

TEXT
User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Meta-WebIndexer
Allow: /

Meta's user agents are written in lowercase (meta-externalagent), but the robots.txt standard (RFC 9309) requires crawlers to match names case-insensitively, so either spelling works.

A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.

User agent strings

Meta publishes each string with and without the URL part:

TEXT
Meta-ExternalAgent:
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-externalagent/1.1

Meta-ExternalFetcher:
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-externalfetcher/1.1

Meta-WebIndexer:
meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-webindexer/1.1

FacebookExternalHit:
facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)
facebookexternalhit/1.1

How to verify a Meta request

Meta says a crawler comes from Meta when its source IP address is on the list this command returns, and notes that these addresses change often:

TEXT
whois -h whois.radb.net -- '-i origin AS32934' | grep ^route

Meta publishes no JSON list and doesn't document a reverse DNS check. Questions go to [email protected].

Where to see Meta's crawlers

Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Meta's crawlers won't appear there. Your server and CDN logs record every request with its user agent and IP address.

OneLence AI visibility recognizes Meta-ExternalAgent, Meta-ExternalFetcher and Meta-WebIndexer by their user agents when your server reports page requests to OneLence. It shows which pages each one requested and the visitors Meta AI sends you, next to the same view for OpenAI, Anthropic, Perplexity and Google.

Frequently asked questions

What is Meta-ExternalAgent?

Meta-ExternalAgent is one of Meta's web crawlers. Meta says it crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Its user agent starts with meta-externalagent/1.1.

Should I block Meta-ExternalAgent?

Block it if you don't want Meta to crawl your site for uses such as AI model training. A rule for Meta-ExternalAgent doesn't apply to Meta-WebIndexer, which Meta says helps it cite and link to your content in Meta AI's responses, or to FacebookExternalHit, which builds link previews.

Will blocking Meta-ExternalAgent break Facebook link previews?

No. Link previews on Facebook, Instagram and Messenger come from FacebookExternalHit, a separate crawler with its own robots.txt name.

What is Meta-ExternalFetcher?

Meta-ExternalFetcher fetches individual links at a user's request and supports product functions such as helping AI navigate websites to complete tasks for users. Meta says it may bypass robots.txt because a user requested the fetch.

Why is meta-externalagent crawling my site so much?

Meta's AI crawlers are among the busiest: Fastly reported in August 2025 that they generated 52% of the AI crawler traffic it observed. Meta doesn't document Crawl-delay support or any other way to set a crawl rate, so disallow the paths you don't want crawled, or write to Meta at [email protected].

How can I tell whether a request really came from Meta?

Meta says a crawler whose source IP address is on the list returned by the command whois -h whois.radb.net -- '-i origin AS32934' | grep ^route comes from Meta, and notes that these addresses change often. It publishes no JSON list and no reverse DNS check.