AI crawler guide

Google-Extended: what it controls, and how to opt out of AI Overviews

Google-Extended is a robots.txt setting, not a crawler: it decides whether pages Google already crawls may train Gemini models and ground Gemini's answers. It doesn't affect Google Search, AI Overviews or AI Mode. Since August 31, 2026, Search Console has a separate setting for those.

Last checked against official documentation on October 5, 2026.

Google-Extended is the robots.txt name Google gives to one decision: whether the pages it crawls for other purposes may also be used for Gemini. It doesn't crawl anything itself, so it never shows up in your logs.

Google's AI controls at a glance

ControlWhat it changesWhat it doesn't change
Disallow Google-Extended in robots.txtTraining of future Gemini models, grounding in Gemini Apps and Grounding with Google Search on Vertex AI, and training of the models behind Search's generative AI featuresInclusion and ranking in Google Search, which includes AI Overviews and AI Mode
Search generative AI control in Search ConsoleLinks to your site and your content in AI Overviews, AI Mode and generative AI features in DiscoverThe rest of Google Search
Disallow Googlebot in robots.txtAll of Google Search, AI features included, plus Google Images, Google Video and Google News
nosnippet, data-nosnippet, max-snippet, noindexWhat Search shows from the page
Disallow GoogleOther in robots.txtNo specific product, says Google

What Google-Extended controls

Google describes Google-Extended as a standalone product token that sites can use to manage whether content Google crawls from them may be used for two things:

  • Training future generations of the Gemini models that power Gemini Apps and the Vertex AI API for Gemini.
  • Grounding in Gemini Apps and Grounding with Google Search on Vertex AI, which means providing content from the Google Search index to the model at prompt time to make answers more factual and relevant.

Search Console's help adds one more: to limit training of the models used to generate responses in Search's generative AI features, use Google-Extended.

Google says Google-Extended doesn't affect a site's inclusion in Google Search and isn't used as a ranking signal. Blocking it isn't an SEO decision.

Google says AI is built into Search, which is why robots.txt rules for Googlebot are the control for how sites are crawled for Search. To be eligible as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to appear in Google Search with a snippet, and Google says there are no additional requirements or special optimizations.

So blocking Google-Extended doesn't keep your pages out of AI Overviews or AI Mode.

The Search generative AI control

Search Console now has a separate setting for that. Google began testing it on June 3, 2026, with a subset of site owners in the UK, and says that as of August 31, 2026, it has rolled it out to all websites worldwide.

When you exclude your site with the control:

  • Links to your site and your site's content won't appear in AI Overviews, AI Mode or generative AI features in Discover.
  • Content crawled from your site won't be used as an input to generate an AI response or preview in those features.
  • The rest of Google Search is unaffected. Google says the control isn't used as a ranking or inclusion signal for other parts of Search.

Google says content is excluded within 1 to 2 days after the control goes live, though caching can make some content take longer. To remove your content from Google Search completely, Google points to noindex instead. Google doesn't say whether the control affects the Gemini app, so Google-Extended remains the documented setting for Gemini Apps.

robots.txt rules for Google-Extended

Opt out of Gemini training and grounding while staying in Google Search:

TEXT
User-agent: Google-Extended
Disallow: /

Opt out for part of your site, following Google's own example:

TEXT
User-agent: Google-Extended
Allow: /archive/1Q84
Disallow: /archive/

Google's crawlers always obey robots.txt rules when they crawl automatically. Google doesn't support Crawl-delay.

A crawler that finds a group with its own name follows only that group and ignores the User-agent: * group. If your * group disallows paths such as /admin/, repeat those lines in each named group that should still respect them.

GoogleOther and Google-Agent

GoogleOther is a generic crawler that Google's product teams may use to fetch publicly accessible content, for example for one-off crawls for internal research and development. Google says robots.txt rules for GoogleOther don't affect any specific product.

Google-Agent is different: agents hosted on Google infrastructure use it to navigate the web and perform actions when a user asks. Because a user requested the fetch, Google says these fetchers generally ignore robots.txt rules.

How to verify Google's crawlers

Google-Extended sends no requests, so there's nothing to verify. For the crawlers that do the crawling, Googlebot and GoogleOther among them, check the source IP address against Google's common-crawlers.json, or run a reverse DNS lookup: the hostname should be in googlebot.com, google.com or googleusercontent.com, and a forward lookup of that hostname should return the original IP address. Google-Agent uses the addresses in user-triggered-agents.json.

Where to see Google's crawlers

Google Analytics 4 automatically excludes traffic from known bots and spiders, so Googlebot won't appear there. Your server and CDN logs record its requests. Google-Extended won't appear anywhere, since no request carries it.

OneLence AI visibility counts Googlebot as classic search rather than AI. It shows the visitors Gemini sends you and the pages Google's user-triggered agents, such as Google-Agent, fetch, confirmed against Google's IP lists, next to the same view for ChatGPT, Claude and Perplexity.

Frequently asked questions

What is Google-Extended?

A robots.txt product token that lets sites choose whether content Google crawls from them may be used to train future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It has no user agent of its own: Google crawls with its existing user agents.

Why don't I see Google-Extended in my server logs?

Because it isn't a crawler. Google says Google-Extended has no separate HTTP request user agent string and that crawling is done with Google's existing user agents. The token only controls how Google may use what it crawls.

Does blocking Google-Extended hurt my SEO?

No. Google says Google-Extended doesn't affect a site's inclusion in Google Search and isn't used as a ranking signal in Google Search.

Does blocking Google-Extended remove my site from AI Overviews?

No. AI Overviews and AI Mode are part of Google Search, which Google-Extended doesn't affect. To opt out of them, use the Search generative AI control in Search Console, which Google rolled out to all websites worldwide as of August 31, 2026. It removes links to your site and your content from AI Overviews, AI Mode and generative AI features in Discover without affecting the rest of Search.

How do I opt out of AI Overviews and AI Mode?

Use the Search generative AI control in Search Console. Google says content is excluded within 1 to 2 days after the control goes live, though caching can make some content take longer, and that the control isn't used as a ranking or inclusion signal for the rest of Search. To limit what Search shows from a page, Google also supports nosnippet, data-nosnippet, max-snippet and noindex.

What is GoogleOther?

A generic Google crawler that product teams may use to fetch publicly accessible content, for example for one-off crawls for internal research and development. Google says robots.txt rules for GoogleOther don't affect any specific product.