Free tool

Robots.txt checker and validator for AI crawlers

See whether GPTBot, ClaudeBot, PerplexityBot, Googlebot and nine other crawlers may read a page, which line of robots.txt decides it, whether your llms.txt follows the format and how much of the page a crawler gets without JavaScript.

Free, no signup. The checker fetches the page, robots.txt and llms.txt once, the way a crawler would.

What the checker looks at

The checker fetches four files once each, the way a crawler would: the page you enter, then robots.txt, llms.txt and llms-full.txt from the site the page ends up on after redirects. It identifies itself as OneLence-Checker, follows up to five redirects and waits up to eight seconds for each file.

  • robots.txt. Whether each of 13 crawlers may read the page, and the line that decides it. The rules are applied the way Google's open-source robots.txt parser applies them, including the common typos it accepts.
  • llms.txt. Whether the site has one and whether it follows the format at llmstxt.org. More in the llms.txt guide.
  • The page without JavaScript. The HTTP status, title, canonical URL, noindex and nosnippet settings, structured data and how much text is in the HTML. OpenAI, Anthropic and Perplexity don't say whether their crawlers run JavaScript, so text that only appears after scripts run may be invisible to them.
  • Cloudflare. Whether the page comes back as a challenge instead of content. Cloudflare can also block AI crawlers before they read robots.txt, which no outside check can see, so the checker says when a site runs on Cloudflare.

How robots.txt decides who may read a page

  1. A crawler follows the group that names it. User-agent lines are matched without regard to case, and when several groups name the same crawler, their rules are combined.
  2. A crawler with its own group ignores the * group. Once robots.txt has a group for GPTBot, GPTBot skips everything under User-agent: *, even if its own group is empty. Only crawlers that no group names follow the * group, and Apple's Applebot first falls back to the rules for Googlebot.
  3. The longest matching rule wins. Within the group, the Allow or Disallow rule with the longest matching path decides. When an Allow and a Disallow rule are equally long, Allow wins, and a page no rule matches is allowed.
  4. * matches anything and $ marks the end. Disallow: /*.pdf$ blocks every URL that ends in .pdf. Paths are case-sensitive, so Disallow: /Blog/ doesn't block /blog/.
  5. Errors count too. A robots.txt that answers 404 or another 4xx error means there are no rules. A 5xx error, a 429 or no answer means crawlers that follow the standard stay away from the whole site until it loads again.

The crawlers it checks

The same 13 crawlers OneLence's own site check uses, grouped by what they do. The AI crawler guide explains each one and lists the IP ranges that verify them.

CrawlerCompanyWhat it doesFollows robots.txt
ChatGPT-UserOpenAIOpens pages when someone asks ChatGPTNot always, says OpenAI
Claude-UserAnthropicOpens pages when someone asks ClaudeYes
Perplexity-UserPerplexityOpens pages when someone asks PerplexityNot always, says Perplexity
OAI-SearchBotOpenAIChatGPT searchYes
Claude-SearchBotAnthropicClaude's web searchYes
PerplexityBotPerplexityPerplexity searchYes
GooglebotGoogleGoogle Search, including AI Overviews and AI ModeYes
BingbotMicrosoftBing searchYes
ApplebotAppleSearch in Siri, Spotlight and SafariYes
GPTBotOpenAITraining OpenAI's modelsYes
ClaudeBotAnthropicTraining Anthropic's modelsYes
Google-ExtendedGoogleGemini training and grounding, without affecting Google SearchA robots.txt name, not a crawler
Applebot-ExtendedAppleTraining Apple's models, without affecting Apple's searchA robots.txt name, not a crawler

Frequently asked questions

How do I check if my robots.txt blocks AI crawlers?

Enter a domain or page URL in the checker above. It reads your robots.txt the way Google's open-source parser does and shows, for 13 crawlers from OpenAI, Anthropic, Perplexity, Google, Microsoft and Apple, whether each may read the page and which line decides it. To test rules before you publish them, paste them in instead.

How do I validate my robots.txt?

Enter your domain or paste your rules into the checker. Besides each crawler's verdict, it flags the lines crawlers skip or read differently: lines that aren't robots.txt rules, misspelled directives, rules that come before any User-agent line, names no crawler matches, paths that don't start with / or *, and Noindex rules, which Google stopped supporting in 2019.

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot only collects training data. ChatGPT search relies on OAI-SearchBot, and ChatGPT-User opens pages while ChatGPT answers someone. Each has its own robots.txt name, and OpenAI says each setting is independent, so a rule for GPTBot doesn't apply to the other two.

Why does a crawler ignore my User-agent: * rules?

Because robots.txt has a group that names it. A crawler follows only the group for its own name, so once there's a User-agent: GPTBot group, GPTBot skips everything under User-agent: *, even if its own group is empty. Repeat the rules you want it to keep in its own group.

Is robots.txt enough to keep AI crawlers out?

No. robots.txt is a request, not a lock. OpenAI says its rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and anyone can send requests under a crawler's name. To stop requests outright, block them at your server or CDN, and check crawlers against the IP lists their operators publish.

What happens if robots.txt is missing or returns an error?

A 404 or another 4xx error means there are no rules, so crawlers may read everything. A 5xx error, a 429 or no answer means the opposite: crawlers that follow the standard stay away from the whole site. Google stops crawling for up to 12 hours, then uses the last copy of robots.txt it saved.

Do I need an llms.txt file?

No, it's optional. llms.txt is a proposed markdown file that points AI agents to a site's most useful pages. Google says Google Search doesn't use it, and OpenAI, Anthropic and Perplexity don't say their crawlers read it on other sites. If you have one, the checker validates it against the format at llmstxt.org: one H1 title, H2 sections and a link list in each.

Why can't AI assistants read my page when robots.txt allows them?

The usual reasons are a firewall or CDN that blocks them before they reach the page, such as Cloudflare's bot settings, and pages that only show their text after JavaScript runs. OpenAI, Anthropic and Perplexity don't say whether their crawlers run JavaScript, so a page built in the browser may look empty to them. The checker shows how much text your HTML holds without scripts.