Free tool

Free robots.txt generator

Choose which search engines and AI crawlers may read your site, keep private paths out and add your sitemap. The file updates as you choose, ready to copy or download, with setup steps for WordPress, Shopify, Webflow and Wix.

Start from

For your sitemap's address and to check the live file afterwards.

Your robots.txt

# Made with the free robots.txt generator at https://onelence.com/ai-crawlers/robots-txt-generator

# Every crawler
User-agent: *
Allow: /

What this file does

  • Model trainingAll 9 allowed
  • AI searchAll 6 allowed
  • Answering peopleAll 8 allowed
  • Search enginesAll 3 allowed
  • Crawlers not listedAllowed

Add your sitemap below so crawlers find all your pages.

For /:

Allowed (26)GPTBotClaudeBotGoogle-ExtendedApplebot-ExtendedMeta-ExternalAgentAmazonbotCCBotMistralAI-TrainingBytespiderOAI-SearchBotClaude-SearchBotPerplexityBotMeta-WebIndexerAmzn-SearchBotMistralAI-IndexChatGPT-UserClaude-UserPerplexity-UserGoogle-AgentMeta-ExternalFetcherAmzn-UserDuckAssistBotMistralAI-UserGooglebotBingbotApplebot

Crawlers not listedAllowed by Allow: /, line 5

Upload it so it answers at https://yourdomain.com/robots.txt. How to add it on WordPress, Shopify, Webflow and Wix. Then check the live file with the robots.txt checker.

Crawlers

Each of these reads its own robots.txt group, so allowing or blocking one leaves the others as they are.

Crawlers not listed here

SEO tools, archives and every other bot, through the User-agent: * group.

Model training

Collect pages that may be used to train AI models. Blocking them doesn't remove you from ChatGPT search, Claude's search or Google Search.

  • GPTBotOpenAI

    Training OpenAI's models · Guide

  • ClaudeBotAnthropic

    Training Anthropic's models · Guide

  • Google-ExtendedGoogle

    Gemini training and grounding, without affecting Google Search · Guide

    A robots.txt name, not a crawler: Google reads it when its other crawlers fetch your pages.

  • Applebot-ExtendedApple

    Training Apple's models, without affecting Apple's search · Guide

    A robots.txt name, not a crawler: Apple reads it for pages Applebot fetches.

  • Meta-ExternalAgentMeta

    Uses such as training Meta's AI models · Guide

  • AmazonbotAmazon

    Improving Amazon's products, and may train Amazon's AI models · Guide

  • CCBotCommon Crawl

    An open web archive that is widely used to train AI models · Guide

  • MistralAI-TrainingMistral

    Training Mistral's models

  • BytespiderByteDance

    Not documented by ByteDance · Guide

    Independent reports say Bytespider doesn't reliably follow robots.txt.

AI search

Build the indexes AI search answers draw on. Blocking them can keep your pages out of those answers.

  • OAI-SearchBotOpenAI

    ChatGPT search · Guide

  • Claude-SearchBotAnthropic

    Claude's web search · Guide

  • PerplexityBotPerplexity

    Perplexity search · Guide

  • Meta-WebIndexerMeta

    Meta AI search and citations · Guide

  • Amzn-SearchBotAmazon

    Search experiences such as Alexa, not training · Guide

  • MistralAI-IndexMistral

    Mistral's search, not training

    Mistral doesn't say whether MistralAI-Index follows robots.txt.

Answering people

Fetch a page while an assistant answers someone. Several may not follow robots.txt, say the companies that run them.

  • ChatGPT-UserOpenAI

    Opens pages when someone asks ChatGPT · Guide

    OpenAI says robots.txt rules may not apply to ChatGPT-User.

  • Claude-UserAnthropic

    Opens pages when someone asks Claude · Guide

  • Perplexity-UserPerplexity

    Opens pages when someone asks Perplexity · Guide

    Perplexity says Perplexity-User generally ignores robots.txt.

  • Google-AgentGoogle

    Agents acting on a person's request

    Google says its user-triggered agents generally ignore robots.txt.

  • Meta-ExternalFetcherMeta

    Fetches links for people and AI agent tasks · Guide

    Meta says Meta-ExternalFetcher may bypass robots.txt.

  • Amzn-UserAmazon

    Live requests, such as Alexa questions · Guide

    Amazon says Amzn-User may not follow every directive.

  • DuckAssistBotDuckDuckGo

    Answers in DuckAssist, not training

    DuckDuckGo says robots.txt changes apply within 72 hours.

  • MistralAI-UserMistral

    Answers in Vibe, formerly Le Chat

Search engines

Search results. What Googlebot may read also decides what AI Overviews and AI Mode can use.

  • GooglebotGoogle

    Google Search, including AI Overviews and AI Mode

  • BingbotMicrosoft

    Bing search

  • ApplebotApple

    Search in Siri, Spotlight and Safari · Guide

    With no group for Applebot, Apple follows the rules for Googlebot.

Private paths

Every crawler you allow skips these. A path covers every address that starts with it, so /admin/ also covers /admin/users. Use * for any characters and $ for the end, as in /*.pdf$. Allow lines make exceptions.

Sitemaps

Use the full address. WordPress serves /wp-sitemap.xml, Yoast SEO and Rank Math /sitemap_index.xml, and Shopify, Wix, Webflow and Squarespace /sitemap.xml.

robots.txt decides who may read your site. OneLence shows who actually does.

OneLence helps marketing teams find growth opportunities, know what to do next, and grow more efficiently without simply spending more. Its AI visibility shows which of your pages AI crawlers read, how many visitors AI assistants send you and which of them become customers. It's in beta on the Premium and Scale plans, and you can try it in the 7-day free trial.

How to use the robots.txt generator

  1. Pick a starting point. Allow every crawler, opt out of AI training, block AI crawlers but keep search engines, or block everything on a staging site.
  2. Adjust crawler by crawler. The list covers 26 crawlers from OpenAI, Anthropic, Google, Perplexity, Microsoft, Apple, Meta, Amazon, Mistral, DuckDuckGo, Common Crawl and ByteDance, grouped by what they do. Crawlers it doesn't list follow the setting at the top.
  3. Add private paths. Admin areas, carts, account pages or site search results. Every crawler you allow skips them, and Allow lines make exceptions.
  4. Add your sitemap. Crawlers read the Sitemap line whichever group they follow, so it goes at the end of the file.
  5. Copy or download the file and upload it. Then enter your domain in the robots.txt checker to see what each crawler finds on your live site.

The generator reads the file back with the checker's parser, which follows Google's open-source robots.txt parser, so "What this file does" and "Test a path" show what a crawler that follows the standard will do with it.

Which AI crawlers to allow or block

AI companies run separate crawlers for separate jobs, each with its own robots.txt name, so you can decide job by job.

  • Model training. GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot and others collect pages that may be used to train AI models. Blocking them doesn't remove you from ChatGPT search, Claude's search or Google Search. Google-Extended and Applebot-Extended aren't crawlers: they tell Google and Apple whether pages their other crawlers fetch may train their models, and Google-Extended also covers grounding in Gemini Apps and Vertex AI.
  • AI search. OAI-SearchBot, Claude-SearchBot and PerplexityBot build the indexes behind ChatGPT search, Claude's search and Perplexity. OpenAI says sites that block OAI-SearchBot aren't shown in ChatGPT search answers, though they can still appear as navigational links, so keep these allowed if you want AI search to cite you.
  • Answering people. ChatGPT-User, Claude-User, Perplexity-User and others open a page when someone's question needs it. OpenAI, Perplexity, Google, Meta and Amazon say their agents may not follow robots.txt, so keeping them out takes a firewall rule. Anthropic says its agents, Claude-User included, follow it.
  • Search engines. Googlebot, Bingbot and Applebot. AI Overviews and AI Mode are part of Google Search, so they follow Googlebot's rules. Since August 31, 2026, a Search Console setting can opt a site out of them without leaving Search.

The AI crawler guide explains each crawler, the IP lists that verify them and what blocking each one changes.

Example robots.txt files

Opt out of AI training

Asks the nine training crawlers to stay away and leaves AI search, AI answers and search engines alone. It's the "Opt out of AI training" starting point above.

# Model training
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: MistralAI-Training
Disallow: /

User-agent: Bytespider
Disallow: /

# Every other crawler
User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

WordPress

The rules WordPress serves on its own, with its built-in sitemap. admin-ajax.php stays open because themes and plugins load content on public pages through it.

# Every crawler
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml

Allow only Google and Bing

Every other crawler is blocked. Applebot follows Googlebot's rules when no group names it, so it gets a group of its own to stay blocked.

# Search engines
User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

User-agent: Applebot
Disallow: /

# Every other crawler
User-agent: *
Disallow: /

Block every crawler

For a staging or test site. Pages already in search results can stay listed, so use a password to keep a site private.

# Every crawler
User-agent: *
Disallow: /

How to add robots.txt to your site

Crawlers only look for the file at the root of each host, so it has to answer at yourdomain.com/robots.txt, and a subdomain such as shop.yourdomain.com needs its own. These steps were checked in October 2026.

  • Any host or static site. Upload robots.txt to your site's root folder. In Next.js, Nuxt and Astro that's the public folder, and in Hugo and Gatsby it's static.
  • WordPress. WordPress serves a robots.txt of its own until a real file replaces it, so you can upload yours to the site's root folder with your host's file manager or FTP. In Yoast SEO, go to Yoast SEO, Tools, File editor, and create or edit the file there; the option only appears when the site allows file editing. In Rank Math, switch to Advanced Mode, then open General Settings, Edit robots.txt, after deleting any robots.txt file in the root folder, since the editor only works without one.
  • Shopify. Choose Shopify above the file to get a robots.txt.liquid template, and save a copy of your current /robots.txt. Then go to Online Store, Themes, click the three dots and Edit code, add a new file in the Templates folder named robots.txt.liquid and paste the template. It keeps Shopify's own rules and sitemap and adds yours. Shopify says the template's default rules don't always match the file it generates, so compare the two after you save.
  • Webflow. In Site settings, SEO, Indexing, paste the rules into the robots.txt field, save, then publish the site. Webflow adds your sitemap's address on its own.
  • Wix. In your dashboard, open SEO & GEO, then Robots.txt Editor under Tools and settings. Click View File, paste the rules, click Save Changes, then Save. Reset to Default brings back Wix's own file.
  • Squarespace. Squarespace serves the same robots.txt for every site and doesn't let you edit it. To ask AI crawlers to stay away, turn on "Block known artificial intelligence crawlers" in Settings, Crawlers. Squarespace notes this may make your site less visible in AI chatbots.
  • Cloudflare. On any plan, Cloudflare's managed robots.txt can add rules that block AI training crawlers ahead of your own, or serve them when you have no file. Its lines combine with yours, so check the result with the robots.txt checker, which shows the lines Cloudflare added.

Rules every robots.txt follows

  • One file per host, at the root. It's named robots.txt in lower case and is plain UTF-8 text. Google reads the first 500 KiB and ignores the rest.
  • A crawler follows the group that names it. Once a crawler has a group of its own, it ignores the User-agent: * group, so private paths have to be repeated in each named group that lets a crawler in. The generator repeats them for you.
  • The longest matching path wins. When an Allow and a Disallow rule are equally long, Allow wins. Paths are case-sensitive, * matches any characters and $ marks the end.
  • Errors count. A robots.txt that answers 404 means there are no rules. A server error makes Google stop crawling for 12 hours, then use its last copy of the file for up to 30 days.
  • It controls crawling, not indexing. A blocked page can still appear in search results, without a description, when other pages link to it. A noindex tag on a page crawlers may read, or a password, keeps it out.
  • Crawl-delay is left out. Google ignores it, while Bing, Anthropic's crawlers and CCBot follow it. Add it by hand if one of them loads your server too much.

Frequently asked questions

What does a robots.txt generator do?

It writes the rules for you: which crawlers may read your site, which paths they should skip and where your sitemap is. This one lists 26 search engine and AI crawlers by what they do, writes the file as you choose and reads it back with the same parser as the robots.txt checker, so you can see what each crawler will do before you upload it.

Is the robots.txt generator free?

Yes. There's no signup, the file is built in your browser, and it's yours to use however you like, on as many sites as you want.

Should I block AI crawlers in robots.txt?

It depends on what you want from AI. To keep your pages out of model training but stay in AI answers, block the training crawlers, such as GPTBot, ClaudeBot and Google-Extended, and leave the AI search and answer agents alone. Blocking AI search crawlers such as OAI-SearchBot or PerplexityBot can keep your pages out of ChatGPT search and Perplexity, and with them the visitors those answers send.

Will blocking AI crawlers hurt my Google rankings?

No. Google Search relies on Googlebot, and Google says blocking Google-Extended doesn't affect inclusion or ranking in Google Search. AI Overviews and AI Mode are part of Search, so they follow Googlebot's rules. Since August 31, 2026, a Search Console setting can opt a site out of them without leaving Search.

Where do I put the robots.txt file?

At the root of your domain, so it answers at yourdomain.com/robots.txt. Crawlers only look there, and each subdomain needs its own file. WordPress SEO plugins, Shopify, Webflow and Wix each have their own place for it, and the steps for each are on this page. Squarespace doesn't let sites edit robots.txt.

How do I allow only Google and Bing?

Choose "Block every crawler", then set Googlebot and Bingbot to Allow. The generator gives each of them a group of its own and blocks everything else. Apple's Applebot follows Googlebot's rules when no group names it, so the generator also writes a group that keeps Applebot out, unless you allow it too.

Does robots.txt remove a page from Google?

No. robots.txt stops crawling, not indexing, so Google can still list a blocked page without a description when other pages link to it. To keep a page out of search results, add a noindex tag and let Google crawl the page so it sees the tag, or put the page behind a password.

How long until crawlers follow a new robots.txt?

Usually about a day. Google caches robots.txt for up to 24 hours, and OpenAI says changes take about 24 hours to reach ChatGPT search. DuckDuckGo says DuckAssistBot applies changes within 72 hours. In Search Console's robots.txt report, you can ask Google to fetch the new file sooner.