Skip to content

How to block AI crawlers on WordPress, or let them in: robots.txt, Cloudflare, hosts and plugins

AI companies send different crawlers for training, for search and for fetching a page someone asked about. A WordPress site can refuse one kind and admit the others in robots.txt. A CDN, a host or a plugin can also block them where robots.txt does not show it.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • Training, search and fetching a page for a user are done under different crawler names. Refusing one does not refuse the others.
  • A fresh WordPress names no AI crawler in robots.txt, so each one follows the same three lines as every other crawler.
  • A crawler that robots.txt names follows only its own group. A group written for it has to repeat any rule it should still obey.
  • robots.txt is a request. OpenAI, Perplexity and Google each say a fetch that a user asked for may not follow it.
  • Cloudflare, a host or a plugin can refuse AI crawlers, and the WordPress dashboard does not show it.
  • A block aimed at a crawler cannot be seen by loading the page yourself. The server's access log shows what each crawler was given.

There is no single "AI crawler" to block or allow. The companies run separate crawlers for separate jobs. OpenAI and Anthropic each have one that collects pages that may be used to train models, one that builds a search index for the assistant's answers, and one that fetches a single page because a person asked about it. Perplexity has the last two. Each answers to its own name in robots.txt, so a WordPress site can refuse the training crawlers and admit the search ones, or the reverse.

A fresh WordPress does neither. Its robots.txt names no AI crawler at all, which leaves every one of them free to crawl. If a site is blocked anyway, the block is in one of three places an owner does not usually look: a line in robots.txt that something else added, a setting at a CDN or a host that stands in front of WordPress, or a plugin.

The crawlers, and what each one is for

Each row is what the company's own page says, read on October 8, 2026. The name is the one to write after User-agent: in robots.txt.

CompanyName in robots.txtWhat the company says it is forWhat it says about robots.txt
OpenAIGPTBotCrawling content that may be used to train its modelsDisallowing it signals that the site's content should not be used for training
OpenAIOAI-SearchBotShowing sites in ChatGPT's search featuresThe name to use for opting out of search
OpenAIChatGPT-UserVisiting a page when a user of ChatGPT or a custom GPT asks. Not automatic crawling, and not what decides searchMay not apply, because a user started the fetch (quoted below)
OpenAIOAI-AdsBotChecking pages submitted as ads on ChatGPT. Not used for trainingThe page does not say
AnthropicClaudeBotCollecting content that could go toward training its modelsHonored, as for all three of its bots
AnthropicClaude-SearchBotAnalyzing content to improve the quality of search resultsHonored
AnthropicClaude-UserFetching a page when a person asks Claude a questionHonored
PerplexityPerplexityBotShowing and linking sites in Perplexity's search results. Not for training foundation modelsA name site owners can use to manage it
PerplexityPerplexity-UserVisiting a page when a user asks a question. Not for crawling or trainingGenerally ignored, because a user asked for the fetch
GoogleGooglebotGoogle Search, which includes AI Overviews and AI ModeThe control for how a site is crawled for Search
GoogleGoogle-ExtendedNot a crawler. A name that governs whether content Google has crawled may be used to train future Gemini models and to ground answers in Gemini Apps and Vertex AIIt sends no requests of its own. The name works only as a rule in robots.txt
MicrosoftBingbotBing's index, which Copilot's answers draw onBing's guidelines say robots.txt controls crawling, not indexing
Common CrawlCCBotAn open repository of crawled pages that anyone can useChecked before every fetch

Anthropic's page is dated April 7, 2026. Google's page on its common crawlers was last updated on July 14, 2026, and its page on AI features on December 10, 2025. The pages of OpenAI, Perplexity, Bing and Common Crawl carry no date.

Three details in that table decide most of what follows.

Training and search are separate at OpenAI and Anthropic, and Perplexity says neither of its crawlers is for training. OpenAI's page says each setting is independent of the others, and gives the example this page is about: allow OAI-SearchBot to appear in search results while disallowing GPTBot.

Google does not split them the same way. Google's page on AI features says AI is built into Search, and that robots.txt rules for Googlebot are the control for Search as a whole. There is no crawler name that admits Google Search and refuses AI Overviews. Google-Extended is about Google's other products, and Google's page on its crawlers says of it:

Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.

The switch for AI features inside Google Search is a setting in Search Console, not a crawler name. How to appear in Google AI Overviews covers it.

A fetch that a user asked for is the weak point of robots.txt. Three companies say so about their own fetchers. OpenAI's page says of ChatGPT-User:

Because these actions are initiated by a user, robots.txt rules may not apply.

Perplexity's page says Perplexity-User generally ignores robots.txt rules, for the same reason. Google's page on user-triggered fetchers, last updated on August 19, 2026, says those fetchers generally ignore robots.txt too; its list includes Google-Agent, used by agents on Google's infrastructure that act on a user's request. Anthropic is the exception: its page says its bots, Claude-User among them, honor robots.txt.

Common Crawl is a non-profit that publishes what it crawls for anyone to use. A post on its own blog, dated February 13, 2024, says many AI models are trained on text of which Common Crawl's data is a large portion. The same post says that blocking CCBot also keeps a site out of the search indexes built from that data.

How a rule for one crawler works in robots.txt

A robots.txt file is made of groups. A group is one or more User-agent lines followed by the rules for the crawlers those lines name. Two rules of the standard, RFC 9309, matter for AI crawlers, and Google's documentation states both:

  • A crawler follows the group that names it, and only that group. It falls back to the * group only when no group names it. Google's documentation adds that a named group and the * group are not combined.
  • Several User-agent lines can share one set of rules. Four names above one Disallow line is one group that applies to all four.

WordPress writes only a * group: three lines that keep crawlers out of /wp-admin/, and a Sitemap line. Every crawler in the table that honors robots.txt therefore follows those three lines and may fetch everything else. The moment a group names a crawler, that crawler stops reading the * group, so its own group has to repeat the /wp-admin/ lines if they should still apply to it.

WordPress robots.txt covers where the file is, what WordPress answers with, and the three ways to change it. The two examples below use one of them: a small must-use plugin on the robots_txt filter, which keeps WordPress's own lines and adds to them.

Two sets of rules you can use

Both examples are one file, saved as ai-crawler-rules.php in wp-content/mu-plugins/. Create the mu-plugins folder if it is not there; a fresh WordPress does not have it. WordPress loads every PHP file in that folder on each request, with nothing to activate. Use one example or the other, not both: they share a file name so that one replaces the other.

Keep training crawlers out, and let search and user crawlers in

This adds one group. It names OpenAI's and Anthropic's training crawlers, Common Crawl's crawler and Google's name for Gemini, and asks all four to stay out of the whole site.

wp-content/mu-plugins/ai-crawler-rules.php
<?php
/**
 * Plugin Name: AI crawler rules
 * Description: Asks the crawlers that collect for AI training to stay out of the site.
 */

add_filter(
	'robots_txt',
	function ( $output, $public ) {
		$output .= "\nUser-agent: GPTBot\n";
		$output .= "User-agent: ClaudeBot\n";
		$output .= "User-agent: CCBot\n";
		$output .= "User-agent: Google-Extended\n";
		$output .= "Disallow: /\n";

		return $output;
	},
	10,
	2
);
Asks the training crawlers to stay out. Remove a User-agent line to leave that crawler alone.

With the file in place, the site answers /robots.txt with WordPress's lines and then the new group. Here example.com stands in for the site's address.

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
What the site answers with. The same text, with your own address in the Sitemap line, is a complete robots.txt file if you would rather use a file.

OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot and Bingbot are not named, so they go on following the * group and may crawl the site. Nothing has to be written to let them in.

One line needs a decision. Google says Google-Extended governs grounding in Gemini Apps as well as training, so that line also asks Google not to use the site's pages to ground Gemini's answers. Google offers no name that separates the two. Delete the line if that is not what you want.

Let every crawler in, and say so

A site that wants to be read by all of them needs no extra rule on a fresh WordPress. Naming the crawlers anyway puts the decision on record, and keeps it standing if someone later tightens the * group.

wp-content/mu-plugins/ai-crawler-rules.php
<?php
/**
 * Plugin Name: AI crawler rules
 * Description: Names the AI crawlers and gives them the same rules WordPress gives every crawler.
 */

add_filter(
	'robots_txt',
	function ( $output, $public ) {
		$output .= "\nUser-agent: GPTBot\n";
		$output .= "User-agent: OAI-SearchBot\n";
		$output .= "User-agent: ChatGPT-User\n";
		$output .= "User-agent: ClaudeBot\n";
		$output .= "User-agent: Claude-SearchBot\n";
		$output .= "User-agent: Claude-User\n";
		$output .= "User-agent: PerplexityBot\n";
		$output .= "User-agent: Perplexity-User\n";
		$output .= "User-agent: CCBot\n";
		$output .= "User-agent: Google-Extended\n";
		$output .= "Allow: /\n";
		$output .= "Disallow: /wp-admin/\n";
		$output .= "Allow: /wp-admin/admin-ajax.php\n";

		return $output;
	},
	10,
	2
);
Names each crawler and gives it the same rules WordPress gives every crawler.
robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: CCBot
User-agent: Google-Extended
Allow: /
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
What the site answers with. The same text, with your own address in the Sitemap line, is a complete robots.txt file.

Allow: / and Disallow: /wp-admin/ both match an address inside the dashboard. The standard says the rule with the longer path is the one to use, so the dashboard stays closed and everything else is open.

Googlebot and Bingbot are left out on purpose. They are search engines' crawlers first, and they already follow the * group.

If you use a file

A file named robots.txt in the site's root replaces WordPress's answer entirely, and the filter then does nothing. Either of the two texts above is a complete file: WordPress's three lines, a Sitemap line that must be your site's real sitemap address, and the group. Do not use the plugin file and a real file together.

To undo either example, delete the file you added.

What robots.txt cannot do

It asks. It does not prevent. The standard says its rules are not a form of access authorization. Cloudflare's documentation puts the same thing in practical terms: compliance is voluntary, and the file does not stop a crawler at a technical level.

It does not bind a fetch a user asked for, at OpenAI, Perplexity or Google, as their own pages say above.

It is not read at once. OpenAI's page says its systems can take about 24 hours to adjust to a change in robots.txt, and Perplexity's says up to 24 hours. Google's documentation says it generally caches the file for up to 24 hours.

It says nothing about what stands in front of the site. A firewall decides before WordPress is asked, and a crawler it turns away never sees a page. That is the subject of the next section.

What can block a crawler before WordPress is asked

Cloudflare

A site that uses Cloudflare has settings for AI crawlers that are separate from robots.txt and are enforced at Cloudflare's network. They changed on September 15, 2026, and Cloudflare's own pages do not all describe the same state, so each statement below is given with its page and date.

Three kinds of crawler. Cloudflare's documentation, last updated on July 1, 2026, sorts AI crawlers by what they do:

KindCloudflare's definition
SearchCrawlers that collect or index content to answer questions about it later
TrainingCrawlers that take content to train or fine-tune a model, including crawlers used for both training and search
AgentAutomated activity acting in real time on a person's behalf, such as chat fetch bots and browser-use agents

Cloudflare's post of July 1, 2026, gives ChatGPT-User as an example of the Agent kind.

Four settings. Each kind is set for the whole domain. Cloudflare's post of September 15, 2026, lists the choices:

SettingWhat Cloudflare says it does
AllowCrawlers of that kind are let through, unless another setting or firewall rule blocks them
Disallow AI TrainingOffered for Training only. A no-training preference is published in robots.txt. Crawlers that serve both search and training and that Cloudflare designates as Accountable stay allowed for search. Every other training crawler is blocked, including the training crawlers of Amazon, Anthropic, Meta and OpenAI
Block on pages with adsCrawlers of that kind are blocked only on pages where Cloudflare detects an ad
BlockCrawlers of that kind are blocked everywhere

One consequence is easy to miss. The same post says that from 15 September, "Block" and "Block on pages with ads" under Training also apply to crawlers that do both jobs, and it names Applebot, Bingbot and Googlebot. A site that now chooses "Block" for Training to keep its content out of AI models is, by Cloudflare's account, also turning Google's and Bing's search crawlers away. The setting that refuses training and keeps search is "Disallow AI Training", and the post says domains that had chosen either kind of block for Training before that date were moved to it.

The defaults for a new domain. From September 15, 2026, per the same post, a domain being added to Cloudflare is offered one of two presets, depending on whether the site earns money from advertising. Any of the settings can be changed during setup or later.

Site without adsSite with ads
Preference SyncEnabledEnabled
SearchAllowAllow
TrainingAllowDisallow AI Training
AgentAllowBlock on pages with ads

Preference Sync is the feature that publishes a preference in robots.txt. Cloudflare finds the pages with ads itself. Its post of July 1, 2025, says it scans the HTML of a page for the code of common ad units, and adds what browsers report loading.

Domains that were already on Cloudflare. Cloudflare has said two things, and they do not agree.

  • Its press release of July 1, 2026, said the new defaults would also be applied on 15 September to all existing free customers who had not changed their settings by that date.
  • Its post of September 15, 2026, answers the question of what an existing customer needs to do this way:

Nothing, in almost every case. Your current settings carry over on their own.

That post gives a table for domains that never used the three controls. Where the older "Block AI bots" switch was off, all three are now set to Allow. Where it was on, in either of its two forms, Search is Allow, Training is "Disallow AI Training", and Agent is "Block on pages with ads".

Cloudflare's documentation page, meanwhile, was last updated on July 1, 2026. It still describes 15 September as a date ahead, and gives the new default for Training as a block on pages with ads, where the later post gives "Disallow AI Training". The practical conclusion is not to reason from any of these pages about a particular site. Open the domain and read its three settings.

Lines in robots.txt that Cloudflare writes. Since July 1, 2025, Cloudflare has offered a managed robots.txt. Its documentation, last updated on August 3, 2026, says that when the setting is on and the site already answers /robots.txt, Cloudflare puts its own lines in front of the site's and serves both as one file. The lines sit between two comments, # BEGIN Cloudflare Managed content and # END Cloudflare Managed Content. The example on that page disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. The post of September 15, 2026, says the managed robots.txt is being replaced by a feature it calls Bot Preference Sync, and that customers who had it on are moved over. Either way, the robots.txt a crawler receives can hold groups that neither WordPress nor anyone on the site wrote.

Where to look. In the Cloudflare dashboard, open the domain and go to its Security Settings page. The documentation names the controls "Configure AI bot policies", and says the robots.txt preference is found by filtering that page by "Bot traffic". A separate section, AI Crawl Control, has a Crawlers tab that lists each AI crawler with its allowed and unsuccessful requests, and an Overview tab that shows whether the managed robots.txt is on. The documentation says that on the free plan it recognizes AI crawlers by their user-agent strings.

The host

Some hosts filter AI crawlers on their own servers, and the customer's robots.txt has no say in it. Only a host's own documentation is a reason to say what a host does, so this section names the two whose pages were read. For any other host, ask in writing whether AI crawlers are filtered and which ones.

  • SiteGround. Its help page, last updated on September 4, 2025, says it allows by default the crawlers that major AI providers use in a person's chat session, and blocks by default crawlers intended for AI model training. A customer who wants the training crawlers let in has to ask through the Help Desk. A second page, last updated on September 1, 2025, lists what is allowed, including ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot, PerplexityBot and Perplexity-User. Neither page lists the names it counts as training crawlers.
  • Kinsta. Its documentation, last updated on June 12, 2026, describes a "Block AI crawlers" switch under a site's Bot protection in MyKinsta. It says the switch blocks AI crawlers entirely, verified ones included, and does not affect search engine crawlers such as Googlebot and Bing's. The page does not list which names it counts as AI crawlers.

SiteGround's default is the same split as the first example above, made for the customer. It becomes a problem only when an owner believes robots.txt settles the matter in either direction.

A plugin

SEO plugins. Two document a setting for AI crawlers.

  • Yoast SEO's help page describes "Block unwanted bots" under Settings, then Advanced, then Crawl optimization. It has four toggles: Google AdsBot in the free plugin, and in Yoast SEO Premium also the Google-Extended name, OpenAI's GPTBot and Common Crawl's CCBot. The list has no toggle for ClaudeBot, and none for any search or user crawler. So it cannot block ChatGPT's or Perplexity's search crawler by accident, and it does not cover Anthropic's training crawler.
  • All in One SEO's documentation, last updated on August 14, 2026, describes an "Unwanted Bots" tab in its Crawl Cleanup settings, with a box for each crawler and one for all AI crawlers. It says the feature works by adding rules to the robots.txt that WordPress writes, and that it does not write to a real robots.txt file. The page does not list the crawlers.

Where a plugin works through robots.txt, as All in One SEO says it does, its rules show in robots.txt as served. That is why the first check below is to read it. Yoast's page does not say how its toggles are applied.

Security plugins. A security plugin's rate limiting counts requests, not names. The documentation of one widely used security plugin describes limits on how many pages a crawler may request per minute, and says a crawler over the limit is refused for a time with a 503 status. It also warns that strict settings can block legitimate crawlers. No default that singles out AI crawlers is described there. A limit set low by someone years ago is still worth looking for.

How to check what a crawler gets

Work from the outside in. The first two checks take a minute, and the last is the only one that shows what really happened.

  1. Step 1: Read robots.txt as served

    What counts is what a crawler receives, not what an editor or a settings screen shows. Replace example.com with your site's address.

    Look for a group that names one of the crawlers in the table, and for Disallow: / under User-agent: *, which asks every crawler without a group of its own to stay out.

    bash
    curl -s https://example.com/robots.txt
  2. Step 2: Look for lines nobody on the site wrote

    A group you do not recognize came from somewhere. Lines between # BEGIN Cloudflare Managed content and # END Cloudflare Managed Content are Cloudflare's. A group for GPTBot, CCBot or Google-Extended with no such comment around it may be an SEO plugin's toggle. If a file named robots.txt exists in the site's root, that file is the whole answer.

  3. Step 3: Read the settings at the CDN and the host

    On Cloudflare, read the Search, Training and Agent settings for the domain, as described above. At the host, look for a bot or AI crawler setting in the hosting panel, and ask support if there is none to see.

  4. Step 4: Ask for a page under a crawler's name

    This sends one request for the home page that carries the example user-agent string OpenAI publishes for GPTBot, and prints only the status that comes back.

    If it prints 403 while the same address loads in your browser, a rule that matches the crawler's name is refusing it. SiteGround's help pages show a rule of that kind for .htaccess. If it prints 200, you have learned less than it seems, for the reason given after these steps.

    bash
    curl -s -o /dev/null -w '%{http_code}\n' -A 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot' https://example.com/
  5. Step 5: Count what each crawler was given, in the access log

    The server's access log has one line for every request it answered. Replace /path/to/access.log with the log's path. The server's configuration sets it, so it differs from host to host: ask the host if the hosting panel does not offer the log.

    Each line of the result is a count, a crawler's name and a status, such as 2 GPTBot 403. A 200 means the page was sent. A 403 means the request was refused. A 429 means too many requests in a given time. A 503 means the server could not handle the request for the moment, and it is the status the rate limiting described above answers with.

    bash
    for bot in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User PerplexityBot Perplexity-User CCBot bingbot Googlebot; do
      grep -i "$bot" /path/to/access.log | awk -v bot="$bot" '{ print bot, $9 }'
    done | sort | uniq -c

The command reads the status from the ninth item on each line. That is where it sits in the "combined" format, which Apache's documentation defines and which nginx uses when no other format is set. A log kept in another format has the status somewhere else on the line.

What you cannot see from your own computer

A block that applies only to a crawler cannot be seen by loading the page as yourself. Your browser is not the crawler, so the page loads, and the dashboard shows nothing wrong.

Sending the crawler's name from your own computer, as in the fourth check, finds only a rule that goes by the name alone. A firewall can also go by where a request comes from. OpenAI and Perplexity publish the addresses their crawlers use and recommend allowing them, Perplexity's page tells site owners to match on the name and the address together, and Cloudflare's documentation says it verifies a bot by a signature, a published list of addresses or reverse DNS. A request from your computer that borrows the name is not the crawler to a firewall of that kind. It may be let through where the real crawler is blocked, or the other way round.

That leaves two records of what really happened. The access log shows every request that reached the server. A crawler stopped at a CDN never reaches the server, so the log has no line for it at all, and the CDN's own report is the place to look: at Cloudflare, the Crawlers tab of AI Crawl Control.

Opting out of AI training

What each company documents as the way to do it:

  • OpenAI: disallow GPTBot in robots.txt.
  • Anthropic: disallow ClaudeBot in robots.txt, on every subdomain that should be left out. Its page says blocking its addresses in place of a robots.txt rule may not work, because it stops the bot from reading the file.
  • Google: disallow Google-Extended in robots.txt. That covers Gemini training, and grounding in Gemini Apps and Vertex AI. Google's help page for the Search Console control says that control does not affect training.
  • Microsoft: not a robots.txt name. Bing's post of September 22, 2023, says content on a page tagged noarchive is not used to train Microsoft's models, and that with nocache only the address, title and snippet may be. Those are robots meta tags on the page. The post says a page with either tag still appears in Bing's search results, but there is a cost: Bing's guidelines say noarchive also keeps a page out of Copilot's answers. Cloudflare's post of September 15, 2026, says Microsoft is building a no-training preference in robots.txt, aimed at early 2027; that is Cloudflare's statement, not one found on a Microsoft page.
  • Perplexity: its page says neither of its crawlers collects content for training foundation models, and offers no opt-out for training.
  • Common Crawl: disallow CCBot in robots.txt. Its FAQ also points to an opt-out registry.

None of these pages says that content already collected is removed. Anthropic's says a restriction signals that the site's future materials should be left out of its training data, and Google's speaks of future generations of Gemini models. Treat an opt-out as applying from the day it is read.

What each choice costs

If you refuseWhat its owner says follows
GPTBot or ClaudeBotA signal that the site's content is not to be used for training. OpenAI says the setting is independent of search
Google-ExtendedNo training of future Gemini models on the site, and no grounding of Gemini Apps' answers in it. No effect on inclusion or ranking in Google Search
CCBotThe site is left out of Common Crawl's data, for those who train on it and for those who build search indexes from it
OAI-SearchBotThe site is not shown in ChatGPT's search answers. OpenAI says it can still appear as a navigational link
Claude-SearchBotAnthropic says this may reduce the site's visibility and accuracy in its users' search results
Claude-UserAnthropic's system does not retrieve the site's content in response to a user's question
PerplexityBotPerplexity recommends allowing it for a site to appear in its search results
BingbotBing's guidelines list blocking Bingbot in robots.txt among the things to avoid, and say Bing and Copilot rely on the same crawling and index
GooglebotThe site is not crawled for Google Search, AI features included

Letting a crawler in is not the same as being quoted. It lets the site be read, and none of the pages above says that a site its crawler may read will be shown in an answer. What is documented beyond access is in generative engine optimization. An llms.txt file is a different subject: none of the crawler pages above offers it as a way to admit or refuse a crawler.

When to get help

  • Ask the host whether AI crawlers are filtered on its servers, and where the access log is.
  • Hand it over if robots.txt as served has lines you cannot trace, or the log shows a crawler refused and you cannot find what refuses it. WP Ministry checks crawler access as part of AI search optimization, and a one-time fix covers one issue on one site, starting with a free diagnosis that gives you a written cause and a fixed quote.

Common questions

Does WordPress block AI crawlers by default?

No. The robots.txt a fresh WordPress answers with has one group, for every crawler, and it only keeps them out of /wp-admin/. It names no AI crawler. A block on a WordPress site comes from a rule someone added, a plugin, the host or a CDN.

Does blocking GPTBot take my site out of ChatGPT?

Not by OpenAI's account. Its page says GPTBot is for training and OAI-SearchBot is what decides whether a site is shown in ChatGPT's search answers, and that each setting is independent of the others.

Does blocking Google-Extended affect my place in Google Search?

Google says it does not: the name has no effect on a site's inclusion in Google Search and is not used as a ranking signal. It governs Gemini training, and grounding in Gemini Apps and Vertex AI.

Can robots.txt stop an assistant from opening a page a user asks about?

Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them. Anthropic says Claude-User honors robots.txt. A rule at a firewall is what turns a request away, and it turns away everything it matches.

I changed robots.txt. How long until it counts?

OpenAI says about 24 hours for its systems to adjust, and Perplexity says up to 24 hours. Google says it generally caches robots.txt for up to 24 hours.

My site is on Cloudflare and I never changed anything. Are AI crawlers blocked?

It depends on the domain, and Cloudflare's own statements about domains left on their defaults do not agree with each other. Open the domain's Security settings and read what Search, Training and Agent are set to.

More on this subject

Quick Fix, done for you

Quick Fix is $49. One issue, one site, up to about an hour. No fix, no fee. 30-day warranty. It starts with a free diagnosis.