How the robots.txt checker decides
It applies the rules the way Google documents them and RFC 9309 defines them. The parser is the one LogNorm’s GEO audit uses, so the tool and the audit agree.
- A crawler obeys the group that names its product token, such as User-agent: GPTBot. Only when no group names it does it fall back to User-agent: *. A specific group replaces the * group; it doesn't add to it.
- Within those rules, the longest matching path wins, whatever the order in the file. Disallow: /docs loses to Allow: /docs/public for /docs/public/guide.
- When an Allow and a Disallow match with the same length, Allow wins.
- A * matches any run of characters and a trailing $ anchors the end, so Disallow: /*.pdf$ blocks PDFs but not /guide.pdf?download=1.
- Rules are prefix matches. Disallow: /app also blocks /application, so end directory rules with a slash.
The HTTP status matters too. A 404 or 403 on robots.txt means “no rules”, so everything may be crawled. A 5xx or 429 makes Google treat the whole site as off limits until the file loads again. The checker follows up to five redirects and reports each case.
How to use the robots.txt checker
Check the live file in three steps, or paste a draft to test it before you deploy.
- Enter your domainThe tool fetches /robots.txt from that host over a guarded connection and reads the first 500 KiB, as Google does.
- Add a path to testType a page such as /pricing or /blog/launch. Change it as often as you like: the verdicts update in your browser without fetching again.
- Read each rowEvery crawler shows Allowed or Blocked for the home page and your path, the line that decided it, and whether it used its own group or fell back to *.
Editing your file? Switch to Paste a robots.txt. Nothing leaves your browser, so you can try rules for a staging site or a change you haven’t shipped.
Which AI crawlers the checker covers
The checker covers the two big search engines plus 13 AI agents in three groups: AI search crawlers that decide whether assistants can cite you, user fetchers that open a page when someone asks, and training crawlers.
| Token | Owner | Type | What it does |
|---|---|---|---|
| Googlebot | Search engine | Google Search crawling and indexing | |
| Bingbot | Microsoft | Search engine | Bing Search crawling and indexing |
| OAI-SearchBot | OpenAI | AI search | ChatGPT search results and citations |
| Claude-SearchBot | Anthropic | AI search | Claude search results and citations |
| PerplexityBot | Perplexity | AI search | Perplexity answers and citations |
| ChatGPT-User | OpenAI | AI user fetch | Pages ChatGPT opens when a user asks |
| Claude-User | Anthropic | AI user fetch | Pages Claude opens when a user asks |
| Perplexity-User | Perplexity | AI user fetch | Pages Perplexity opens when a user asks |
| GPTBot | OpenAI | AI training | Training OpenAI models |
| ClaudeBot | Anthropic | AI training | Training Anthropic models |
| Google-Extended | AI training | Gemini training and grounding (not Google Search) | |
| Applebot-Extended | Apple | AI training | Apple Intelligence training |
| CCBot | Common Crawl | AI training | Common Crawl, used by many model trainers |
| Bytespider | ByteDance | AI training | ByteDance / Doubao training |
| meta-externalagent | Meta | AI training | Meta AI training and answers |
Google-Extended and Applebot-Extended aren’t separate crawlers. Googlebot and Applebot do the fetching, and these tokens only say whether the content may be used for AI models. Blocking Google-Extended doesn’t affect Google Search.
What the results mean
Allowed means robots.txt lets that crawler fetch the URL. It doesn’t mean the crawler visits, or that your CDN or firewall lets it through.
- Own group: a group names this crawler, so your * rules don’t apply to it at all.
- Falls back to User-agent: *: no group names it, so it follows your general rules.
- No rule matches: nothing in the applicable group covers that path, which means allowed.
- Lines crawlers will ignore: rules before any User-agent line, typos and unknown fields. They do nothing, which is often not what the author meant.
Content-Signal is a newer line from contentsignals.org that states how your content may be used: search, ai-input (quoting in AI answers) and ai-train. It records your preference; crawlers decide whether to honour it.
How to fix common robots.txt problems
Most broken robots.txt files come from one of five mistakes. Each takes a minute to fix once you see it.
Googlebot is blocked from the whole site
Usually a Disallow: / copied from a staging site. Remove it, or scope it to the paths you meant, then re-check here.
AI search crawlers are blocked by accident
A blanket rule for “AI bots” often catches OAI-SearchBot and PerplexityBot along with training crawlers. Give training crawlers their own group and leave search crawlers allowed if you want to show up in answers. Our guide on how to rank in ChatGPT explains why that access matters.
A specific group drops your general rules
Adding User-agent: Googlebot with one rule means Googlebot ignores everything under *. Repeat the rules you still want inside the specific group.
No sitemap line
Add Sitemap: with the full URL. Any crawler can find your sitemap from there, not only the ones you submitted it to.
robots.txt returns a web page or an error
Single-page apps often answer /robots.txt with their HTML shell. Serve a plain-text file with a 200 status. A server error is worse: Google pauses crawling.
A starting point for a startup that wants search and AI answers but no training:
# Search engines and AI search: welcome User-agent: * Allow: / Disallow: /app/ Content-Signal: search=yes, ai-input=yes, ai-train=no # Training crawlers: opt out User-agent: GPTBot User-agent: CCBot Disallow: / Sitemap: https://acme.com/sitemap.xml
LogNorm's GEO audit tests crawler access at your CDN, llms.txt, structured data and agent readiness, and generates the robots.txt section for you.