Skip to content

Are AI crawlers allowed on my site? A robots.txt check

Give your site's address and see what its robots.txt tells each AI crawler by name: the ones behind ChatGPT's, Claude's and Perplexity's search, the fetches made for their users, and the ones that collect pages for training.

Such as example.com. Our server asks that site for its robots.txt, once, and keeps neither the address nor the file.

What it looks at

Our server asks the site you name for its robots.txt, once, and reads what the file tells each AI crawler by the name that crawler's own company gives it. Give a page's address and the answer is for that page; give the site's and it is for the home page.

The crawlers are kept in three kinds, because a site may want one kind and not another.

  • Search crawlers. OAI-SearchBot, Claude-SearchBot and PerplexityBot build the index their assistant searches when it answers. OpenAI says a site that keeps its crawler out is not shown in ChatGPT's search answers, and the other two say that keeping theirs out can reduce how a site shows in theirs.
  • Fetches for a user. ChatGPT-User, Claude-User and Perplexity-User fetch one page because a person asked about it. OpenAI says robots.txt may not apply to its one, since a person started the request.
  • Training crawlers. GPTBot and ClaudeBot collect pages that may be used to train models. Google-Extended is not a crawler but a name in robots.txt that tells Google whether pages it crawls may be used for Gemini. Google says it does not affect whether a site is included in Google Search.

It also reads the file for Googlebot and Bingbot. Google's AI answers are built on Google Search, and Microsoft's Copilot answers from Bing's index, so for those two the ordinary crawler is the one that counts.

What it cannot tell you

robots.txt is a request, and it is all this check reads. It cannot tell you whether a crawler obeys it. It cannot see a firewall, a CDN or a hosting company that turns a crawler away by name before your site is asked: that shows only to the crawler itself, and our request names itself as ours and never as a crawler. To know what a crawler was really given, look for its name in your server's access log, and at the bot settings of anything that stands in front of the site.

And it cannot tell you whether an assistant will quote your site. Being let in is where that starts, not where it ends, and nobody can promise it. Blocking or allowing AI crawlers on WordPress covers each place a block can sit and how to write the rules.

What happens to the address

The address goes to our server, which makes the one request and sends back the report. Neither the address nor the file is kept or logged. The request names itself as WPMinistry-SiteCheck. Only public web addresses are asked.

Common questions

It says a crawler may ask for my pages, but my site never shows in that assistant. Why?
Being allowed by robots.txt is only the first thing. Something in front of the site may still be turning the crawler away, which this check cannot see. And a crawler that reaches a page does not have to use it: which pages an assistant quotes is its own decision.
My robots.txt names none of these crawlers. What rules do they follow?
The ones written for every crawler, in the group that begins "User-agent: *". On a WordPress site with no robots.txt file of its own, that group keeps crawlers out of the admin and nothing else.
Can I keep my content out of AI training and still be found by AI assistants?
As far as robots.txt goes, yes: the companies use different names for the two jobs, so a file can keep the training crawlers out and leave the search crawlers alone. The robots.txt generator writes that file.
Why is a blocked training crawler not marked as a problem?
Because it is a choice many site owners make on purpose, and it does not affect search. A blocked search crawler is marked as worth a look, because it has a cost you may not have meant to pay.

Want it checked from the inside as well?

Tell us the address and what this showed. We look at the site, its logs and what stands in front of it, and reply with what we found and a fixed quote. The diagnosis is free.

Get a free diagnosis