Free tool · no sign-up

Robots.txt checker for Google and AI crawlers

This free robots.txt checker fetches your file and tells you, for 15 crawlers including Googlebot, GPTBot, ClaudeBot and PerplexityBot, whether each may fetch your home page and any path you add. It quotes the exact line that decides, lists your sitemaps and reads your Content-Signal settings. Use it before and after every robots.txt change.

Enter a domain to see whether Googlebot, Bingbot and 13 AI crawlers may fetch its home page and any path you add. Results show the exact line that decides each one.

How the robots.txt checker decides

It applies the rules the way Google documents them and RFC 9309 defines them. The parser is the one LogNorm’s GEO audit uses, so the tool and the audit agree.

  • A crawler obeys the group that names its product token, such as User-agent: GPTBot. Only when no group names it does it fall back to User-agent: *. A specific group replaces the * group; it doesn't add to it.
  • Within those rules, the longest matching path wins, whatever the order in the file. Disallow: /docs loses to Allow: /docs/public for /docs/public/guide.
  • When an Allow and a Disallow match with the same length, Allow wins.
  • A * matches any run of characters and a trailing $ anchors the end, so Disallow: /*.pdf$ blocks PDFs but not /guide.pdf?download=1.
  • Rules are prefix matches. Disallow: /app also blocks /application, so end directory rules with a slash.

The HTTP status matters too. A 404 or 403 on robots.txt means “no rules”, so everything may be crawled. A 5xx or 429 makes Google treat the whole site as off limits until the file loads again. The checker follows up to five redirects and reports each case.

How to use the robots.txt checker

Check the live file in three steps, or paste a draft to test it before you deploy.

  1. Enter your domain
    The tool fetches /robots.txt from that host over a guarded connection and reads the first 500 KiB, as Google does.
  2. Add a path to test
    Type a page such as /pricing or /blog/launch. Change it as often as you like: the verdicts update in your browser without fetching again.
  3. Read each row
    Every crawler shows Allowed or Blocked for the home page and your path, the line that decided it, and whether it used its own group or fell back to *.

Editing your file? Switch to Paste a robots.txt. Nothing leaves your browser, so you can try rules for a staging site or a change you haven’t shipped.

Which AI crawlers the checker covers

The checker covers the two big search engines plus 13 AI agents in three groups: AI search crawlers that decide whether assistants can cite you, user fetchers that open a page when someone asks, and training crawlers.

Crawlers the robots.txt checker reports
TokenOwnerTypeWhat it does
GooglebotGoogleSearch engineGoogle Search crawling and indexing
BingbotMicrosoftSearch engineBing Search crawling and indexing
OAI-SearchBotOpenAIAI searchChatGPT search results and citations
Claude-SearchBotAnthropicAI searchClaude search results and citations
PerplexityBotPerplexityAI searchPerplexity answers and citations
ChatGPT-UserOpenAIAI user fetchPages ChatGPT opens when a user asks
Claude-UserAnthropicAI user fetchPages Claude opens when a user asks
Perplexity-UserPerplexityAI user fetchPages Perplexity opens when a user asks
GPTBotOpenAIAI trainingTraining OpenAI models
ClaudeBotAnthropicAI trainingTraining Anthropic models
Google-ExtendedGoogleAI trainingGemini training and grounding (not Google Search)
Applebot-ExtendedAppleAI trainingApple Intelligence training
CCBotCommon CrawlAI trainingCommon Crawl, used by many model trainers
BytespiderByteDanceAI trainingByteDance / Doubao training
meta-externalagentMetaAI trainingMeta AI training and answers

Google-Extended and Applebot-Extended aren’t separate crawlers. Googlebot and Applebot do the fetching, and these tokens only say whether the content may be used for AI models. Blocking Google-Extended doesn’t affect Google Search.

What the results mean

Allowed means robots.txt lets that crawler fetch the URL. It doesn’t mean the crawler visits, or that your CDN or firewall lets it through.

  • Own group: a group names this crawler, so your * rules don’t apply to it at all.
  • Falls back to User-agent: *: no group names it, so it follows your general rules.
  • No rule matches: nothing in the applicable group covers that path, which means allowed.
  • Lines crawlers will ignore: rules before any User-agent line, typos and unknown fields. They do nothing, which is often not what the author meant.

Content-Signal is a newer line from contentsignals.org that states how your content may be used: search, ai-input (quoting in AI answers) and ai-train. It records your preference; crawlers decide whether to honour it.

How to fix common robots.txt problems

Most broken robots.txt files come from one of five mistakes. Each takes a minute to fix once you see it.

Googlebot is blocked from the whole site

Usually a Disallow: / copied from a staging site. Remove it, or scope it to the paths you meant, then re-check here.

AI search crawlers are blocked by accident

A blanket rule for “AI bots” often catches OAI-SearchBot and PerplexityBot along with training crawlers. Give training crawlers their own group and leave search crawlers allowed if you want to show up in answers. Our guide on how to rank in ChatGPT explains why that access matters.

A specific group drops your general rules

Adding User-agent: Googlebot with one rule means Googlebot ignores everything under *. Repeat the rules you still want inside the specific group.

No sitemap line

Add Sitemap: with the full URL. Any crawler can find your sitemap from there, not only the ones you submitted it to.

robots.txt returns a web page or an error

Single-page apps often answer /robots.txt with their HTML shell. Serve a plain-text file with a 200 status. A server error is worse: Google pauses crawling.

A starting point for a startup that wants search and AI answers but no training:

# Search engines and AI search: welcome
User-agent: *
Allow: /
Disallow: /app/
Content-Signal: search=yes, ai-input=yes, ai-train=no

# Training crawlers: opt out
User-agent: GPTBot
User-agent: CCBot
Disallow: /

Sitemap: https://acme.com/sitemap.xml
Check what AI crawlers see across your site

LogNorm's GEO audit tests crawler access at your CDN, llms.txt, structured data and agent readiness, and generates the robots.txt section for you.

See the GEO audit

Questions people ask

How do I check if my robots.txt blocks GPTBot?

Enter your domain in the checker above. The GPTBot row shows Allowed or Blocked for your home page and for any path you type, plus the robots.txt line that decided it. If no group names GPTBot, it follows your User-agent: * rules.

Should I block AI crawlers in robots.txt?

Block training crawlers only if you don't want your content used to train models; that is a business choice. Keep AI search crawlers and user fetchers such as OAI-SearchBot, Claude-SearchBot, PerplexityBot and ChatGPT-User allowed if you want AI assistants to cite and link to you. Blocking them removes you from those answers.

What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

They are three OpenAI agents with separate robots.txt tokens. GPTBot collects content for training OpenAI models. OAI-SearchBot finds pages for ChatGPT search results and citations. ChatGPT-User fetches a page when a person asks ChatGPT to open it. You can allow or block each one on its own.

Does Disallow in robots.txt remove a page from Google?

No. Disallow stops Googlebot from crawling the page, but Google can still index the URL if other pages link to it, just without its content. To keep a page out of results, allow crawling and add a noindex robots meta tag or X-Robots-Tag header.

Where does robots.txt have to live?

At the root of each host, for example https://acme.com/robots.txt. A subdomain such as docs.acme.com needs its own file. Crawlers ignore a robots.txt in a subfolder, and Google reads only the first 500 KiB of the file.

Allowed in. Now get cited.

LogNorm's GEO audit checks crawler access at your CDN as well as in robots.txt, then tracks whether ChatGPT, Gemini and Google AI Overviews name you. Your agent ships the fixes.

Already use an agent?

Paste this into Claude Code, Codex, Cursor or Claude. It connects itself; you click Allow once.

Claude CodeClaude CodeCodexCodexCursorCursorClaudeClaudeand any MCP client