How to use the robots.txt generator
- Pick a starting point. Allow every crawler, opt out of AI training, block AI crawlers but keep search engines, or block everything on a staging site.
- Adjust crawler by crawler. The list covers 26 crawlers from OpenAI, Anthropic, Google, Perplexity, Microsoft, Apple, Meta, Amazon, Mistral, DuckDuckGo, Common Crawl and ByteDance, grouped by what they do. Crawlers it doesn't list follow the setting at the top.
- Add private paths. Admin areas, carts, account pages or site search results. Every crawler you allow skips them, and Allow lines make exceptions.
- Add your sitemap. Crawlers read the Sitemap line whichever group they follow, so it goes at the end of the file.
- Copy or download the file and upload it. Then enter your domain in the robots.txt checker to see what each crawler finds on your live site.
The generator reads the file back with the checker's parser, which follows Google's open-source robots.txt parser, so "What this file does" and "Test a path" show what a crawler that follows the standard will do with it.
Which AI crawlers to allow or block
AI companies run separate crawlers for separate jobs, each with its own robots.txt name, so you can decide job by job.
- Model training. GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot and others collect pages that may be used to train AI models. Blocking them doesn't remove you from ChatGPT search, Claude's search or Google Search. Google-Extended and Applebot-Extended aren't crawlers: they tell Google and Apple whether pages their other crawlers fetch may train their models, and Google-Extended also covers grounding in Gemini Apps and Vertex AI.
- AI search. OAI-SearchBot, Claude-SearchBot and PerplexityBot build the indexes behind ChatGPT search, Claude's search and Perplexity. OpenAI says sites that block OAI-SearchBot aren't shown in ChatGPT search answers, though they can still appear as navigational links, so keep these allowed if you want AI search to cite you.
- Answering people. ChatGPT-User, Claude-User, Perplexity-User and others open a page when someone's question needs it. OpenAI, Perplexity, Google, Meta and Amazon say their agents may not follow robots.txt, so keeping them out takes a firewall rule. Anthropic says its agents, Claude-User included, follow it.
- Search engines. Googlebot, Bingbot and Applebot. AI Overviews and AI Mode are part of Google Search, so they follow Googlebot's rules. Since August 31, 2026, a Search Console setting can opt a site out of them without leaving Search.
The AI crawler guide explains each crawler, the IP lists that verify them and what blocking each one changes.
Example robots.txt files
Opt out of AI training
Asks the nine training crawlers to stay away and leaves AI search, AI answers and search engines alone. It's the "Opt out of AI training" starting point above.
# Model training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: MistralAI-Training
Disallow: /
User-agent: Bytespider
Disallow: /
# Every other crawler
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
WordPress
The rules WordPress serves on its own, with its built-in sitemap. admin-ajax.php stays open because themes and plugins load content on public pages through it.
# Every crawler
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xml
Allow only Google and Bing
Every other crawler is blocked. Applebot follows Googlebot's rules when no group names it, so it gets a group of its own to stay blocked.
# Search engines
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /
User-agent: Applebot
Disallow: /
# Every other crawler
User-agent: *
Disallow: /
Block every crawler
For a staging or test site. Pages already in search results can stay listed, so use a password to keep a site private.
# Every crawler
User-agent: *
Disallow: /
How to add robots.txt to your site
Crawlers only look for the file at the root of each host, so it has to answer at yourdomain.com/robots.txt, and a subdomain such as shop.yourdomain.com needs its own. These steps were checked in October 2026.
- Any host or static site. Upload
robots.txtto your site's root folder. In Next.js, Nuxt and Astro that's thepublicfolder, and in Hugo and Gatsby it'sstatic. - WordPress. WordPress serves a robots.txt of its own until a real file replaces it, so you can upload yours to the site's root folder with your host's file manager or FTP. In Yoast SEO, go to Yoast SEO, Tools, File editor, and create or edit the file there; the option only appears when the site allows file editing. In Rank Math, switch to Advanced Mode, then open General Settings, Edit robots.txt, after deleting any robots.txt file in the root folder, since the editor only works without one.
- Shopify. Choose Shopify above the file to get a
robots.txt.liquidtemplate, and save a copy of your current/robots.txt. Then go to Online Store, Themes, click the three dots and Edit code, add a new file in the Templates folder namedrobots.txt.liquidand paste the template. It keeps Shopify's own rules and sitemap and adds yours. Shopify says the template's default rules don't always match the file it generates, so compare the two after you save. - Webflow. In Site settings, SEO, Indexing, paste the rules into the robots.txt field, save, then publish the site. Webflow adds your sitemap's address on its own.
- Wix. In your dashboard, open SEO & GEO, then Robots.txt Editor under Tools and settings. Click View File, paste the rules, click Save Changes, then Save. Reset to Default brings back Wix's own file.
- Squarespace. Squarespace serves the same robots.txt for every site and doesn't let you edit it. To ask AI crawlers to stay away, turn on "Block known artificial intelligence crawlers" in Settings, Crawlers. Squarespace notes this may make your site less visible in AI chatbots.
- Cloudflare. On any plan, Cloudflare's managed robots.txt can add rules that block AI training crawlers ahead of your own, or serve them when you have no file. Its lines combine with yours, so check the result with the robots.txt checker, which shows the lines Cloudflare added.
Rules every robots.txt follows
- One file per host, at the root. It's named
robots.txtin lower case and is plain UTF-8 text. Google reads the first 500 KiB and ignores the rest. - A crawler follows the group that names it. Once a crawler has a group of its own, it ignores the
User-agent: *group, so private paths have to be repeated in each named group that lets a crawler in. The generator repeats them for you. - The longest matching path wins. When an Allow and a Disallow rule are equally long, Allow wins. Paths are case-sensitive,
*matches any characters and$marks the end. - Errors count. A robots.txt that answers 404 means there are no rules. A server error makes Google stop crawling for 12 hours, then use its last copy of the file for up to 30 days.
- It controls crawling, not indexing. A blocked page can still appear in search results, without a description, when other pages link to it. A noindex tag on a page crawlers may read, or a password, keeps it out.
- Crawl-delay is left out. Google ignores it, while Bing, Anthropic's crawlers and CCBot follow it. Add it by hand if one of them loads your server too much.
Frequently asked questions
What does a robots.txt generator do?
It writes the rules for you: which crawlers may read your site, which paths they should skip and where your sitemap is. This one lists 26 search engine and AI crawlers by what they do, writes the file as you choose and reads it back with the same parser as the robots.txt checker, so you can see what each crawler will do before you upload it.
Is the robots.txt generator free?
Yes. There's no signup, the file is built in your browser, and it's yours to use however you like, on as many sites as you want.
Should I block AI crawlers in robots.txt?
It depends on what you want from AI. To keep your pages out of model training but stay in AI answers, block the training crawlers, such as GPTBot, ClaudeBot and Google-Extended, and leave the AI search and answer agents alone. Blocking AI search crawlers such as OAI-SearchBot or PerplexityBot can keep your pages out of ChatGPT search and Perplexity, and with them the visitors those answers send.
Will blocking AI crawlers hurt my Google rankings?
No. Google Search relies on Googlebot, and Google says blocking Google-Extended doesn't affect inclusion or ranking in Google Search. AI Overviews and AI Mode are part of Search, so they follow Googlebot's rules. Since August 31, 2026, a Search Console setting can opt a site out of them without leaving Search.
Where do I put the robots.txt file?
At the root of your domain, so it answers at yourdomain.com/robots.txt. Crawlers only look there, and each subdomain needs its own file. WordPress SEO plugins, Shopify, Webflow and Wix each have their own place for it, and the steps for each are on this page. Squarespace doesn't let sites edit robots.txt.
How do I allow only Google and Bing?
Choose "Block every crawler", then set Googlebot and Bingbot to Allow. The generator gives each of them a group of its own and blocks everything else. Apple's Applebot follows Googlebot's rules when no group names it, so the generator also writes a group that keeps Applebot out, unless you allow it too.
Does robots.txt remove a page from Google?
No. robots.txt stops crawling, not indexing, so Google can still list a blocked page without a description when other pages link to it. To keep a page out of search results, add a noindex tag and let Google crawl the page so it sees the tag, or put the page behind a password.
How long until crawlers follow a new robots.txt?
Usually about a day. Google caches robots.txt for up to 24 hours, and OpenAI says changes take about 24 hours to reach ChatGPT search. DuckDuckGo says DuckAssistBot applies changes within 72 hours. In Search Console's robots.txt report, you can ask Google to fetch the new file sooner.
