Skip to content

How to stop AI training on your WordPress content and stay in AI search answers

You can ask the crawlers that collect for AI training to stay out, one company at a time, and stay in most assistants' search answers. The request is a robots.txt group or, at Microsoft and Amazon, a noarchive tag. It is not a lock, and it does not take back what was already collected.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • OpenAI, Anthropic, Apple, Amazon and Meta give training and search different names in robots.txt, so one can be refused and the other left alone.
  • Google-Extended governs Gemini training and does not affect Google Search. AI Overviews are part of Search and follow Search's own controls.
  • Microsoft has no training name in robots.txt. Its opt-out is a noarchive tag on the page, which also takes the page out of Copilot's answers.
  • Every opt-out is a request about future use. No page cited here says content already collected is taken back out.
  • A noai tag, a line in llms.txt and a notice in the footer are not documented as read by any of these crawlers.
  • The list of names changes. Read each company's page again before relying on any list, this one included.

You can ask the crawlers that collect pages for AI training to stay out of a WordPress site, one company at a time, and each company below that offers an opt-out says it honors it. OpenAI, Anthropic, Apple, Amazon and Meta use one name for training and another for the search behind their assistant, so a site can refuse the first and stay in the second.

Three limits come with that. A request is not a lock. It is about future use: none of the pages cited here says content already collected is taken back out. And at Microsoft the two are tied together, because the tag that refuses training also takes a page out of Copilot's answers.

How to block or allow AI crawlers on WordPress covers how a robots.txt group works and where a CDN, a host or a plugin can block a crawler. This page stays on one question: training out, search answers in.

What each company documents

Each row is from the company's own page, read on October 8, 2026. That is the standing of everything in the table: documented by the company, about its own crawler. Whether a crawler does what its page says is not shown here, and no study of that is cited. A name in the table is one to write after User-agent: in robots.txt, unless the cell says it is a tag.

CompanyWhat governs trainingWhat governs its search answersSeparate?What the company says a block doesIts page
OpenAIGPTBotOAI-SearchBotYes. It says each setting is independentDisallowing GPTBot indicates the content should not be used to train its models. Without OAI-SearchBot a site is not shown in ChatGPT's search answers, though it can still appear as a navigational linkUndated
AnthropicClaudeBotClaude-SearchBotYesRestricting ClaudeBot signals that the site's future materials should be left out of its training data. Disabling Claude-SearchBot may reduce the site's visibility and accuracy in users' search resultsApril 7, 2026
GoogleGoogle-ExtendedGooglebot, and Search's own controlsPartly. See belowGoogle-Extended governs training of future Gemini models, and grounding in Gemini Apps and Vertex AI. It does not affect inclusion or ranking in SearchJuly 14, 2026
MicrosoftA noarchive or nocache robots tag on the page. Its pages give no robots.txt name for trainingBingbot, and the same two tagsNo. One tag governs bothnoarchive: not used for training, and not linked in Copilot. nocache: only the address, title and snippet are used, for bothHelp page undated
AppleApplebot-ExtendedApplebot. For AI answers, a nosnippet tagYesDisallowing Applebot-Extended opts out of training Apple's foundation models. The pages can still be in search results, and the rule is not considered in rankingSeptember 4, 2026
Metameta-externalagentmeta-webindexerTwo names, but the first is described as for training models or indexing content directlyAllowing meta-webindexer helps Meta cite and link the site in Meta AI's responses. The page does not say what refusing meta-externalagent costsMay 21, 2026
AmazonAmazonbot, or a noarchive tag on the pageAmzn-SearchBotYes. It says each setting is independentAmazonbot improves its products and may be used to train its models. noarchive means the page is not used for model training. Amzn-SearchBot makes content eligible for search experiences such as Alexa and does not crawl for trainingUndated
PerplexityNothing. It says neither of its crawlers collects for training foundation modelsPerplexityBotNothing to separateIt recommends allowing PerplexityBot for a site to appear in its search resultsUndated
Common CrawlCCBotIt has no assistant. Others build on its open archiveNo. One crawler, one archiveIts crawler stops crawling the site. Its post of February 13, 2024 says that also keeps the site from those who build search indexes on the dataUndated

Two things in the table are easy to miss.

Two of the names are not crawlers. Google-Extended and Applebot-Extended send no requests. Google and Apple each say the name only decides how pages their ordinary crawler fetched may be used, so neither shows in an access log.

Three of the names do more than training. meta-externalagent, Amazonbot and CCBot each serve other uses too, as their rows say. Refusing one refuses those uses with it.

Google: three controls for three things

  • Training Gemini. Google's page on its crawlers says Google-Extended governs whether content Google crawls may be used to train future generations of Gemini models and to ground answers in Gemini Apps and Vertex AI. It says the name does not affect a site's inclusion in Google Search and is not a ranking signal there. No name covers training alone.
  • AI Overviews and AI Mode. These are part of Search. Google's page on AI features says AI is built into Search, so the rules for Googlebot are the control for how a site is crawled for it. No crawler name admits Search and refuses AI Overviews.
  • How much of a page Search may show or use. Google's specification says nosnippet prevents a page's content from being used as a direct input for AI Overviews and AI Mode, and max-snippet limits how much may be. The cost is in the same sentence: both apply to every kind of result, so the ordinary snippet under the page's title goes or shrinks too. data-nosnippet keeps one part of a page out of snippets.

Search Console also has a setting that takes a whole site out of AI Overviews and AI Mode. Its help page says the setting does not affect AI training, and names Google-Extended as the way to limit training of the models that generate those answers. How to appear in Google AI Overviews covers that setting and the snippet rules.

So at Google, refusing training is one line in robots.txt, and by Google's account it costs nothing in Search.

Microsoft: a tag on the page, not a name in robots.txt

Bing's help page on robots tags gives the two that matter here.

TagIn CopilotFor training Microsoft's models
noarchiveNot linkedNot used
nocacheOnly the address, snippet and title are shownOnly addresses, titles and snippets are used
BothTreated as nocacheTreated as nocache

Bing's post of September 22, 2023, which introduced this, says a page with either tag still appears in Bing's search results. Bing's Webmaster Guidelines state the cost: noarchive prevents content from being used in Copilot's responses, and nocache reduces how much of a page Copilot can cite.

So at Microsoft the two cannot both be had in full. noarchive is the whole opt-out and gives up Copilot. nocache keeps a short form of both.

Two more companies document the same tag. Amazon's page says its crawlers read noarchive as an instruction not to use the page for model training. Google's specification lists noarchive and nocache among rules Google Search does not use and ignores.

The WordPress side

WordPress writes /robots.txt and each page's robots tag itself, and has a filter for each. A small must-use plugin on a filter adds to what WordPress writes and removes nothing.

  1. Step 1: Create the mu-plugins folder if it is not there

    A fresh WordPress has no mu-plugins folder in wp-content, so create it. WordPress loads every PHP file in that folder on each request, with nothing to activate.

  2. Step 2: Add the robots.txt group

    Save this as no-ai-training.php in that folder. It adds one group with the seven training names from the table.

    wp-content/mu-plugins/no-ai-training.php
    <?php
    /**
     * Plugin Name: No AI training
     * Description: Asks the crawlers that collect for AI training to stay out of the site.
     */
    
    add_filter(
    	'robots_txt',
    	function ( $output, $public ) {
    		$output .= "\nUser-agent: GPTBot\n";
    		$output .= "User-agent: ClaudeBot\n";
    		$output .= "User-agent: Google-Extended\n";
    		$output .= "User-agent: Applebot-Extended\n";
    		$output .= "User-agent: meta-externalagent\n";
    		$output .= "User-agent: Amazonbot\n";
    		$output .= "User-agent: CCBot\n";
    		$output .= "Disallow: /\n";
    
    		return $output;
    	},
    	10,
    	2
    );
  3. Step 3: Read robots.txt as served

    Replace example.com with your site's address.

    bash
    curl -s https://example.com/robots.txt
  4. Step 4: Check what came back

    WordPress's own three lines and its Sitemap line are still there, and the new group follows them.

    text
    User-agent: *
    Disallow: /wp-admin/
    Allow: /wp-admin/admin-ajax.php
    
    Sitemap: https://example.com/wp-sitemap.xml
    
    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: Applebot-Extended
    User-agent: meta-externalagent
    User-agent: Amazonbot
    User-agent: CCBot
    Disallow: /
  5. Step 5: Add the tag, if you accept its cost

    Save this as no-ai-training-tag.php in the same folder. It adds noarchive to the robots tag of every page WordPress builds. To keep the short form of Copilot's answers, write nocache in place of noarchive in the line that sets it.

    wp-content/mu-plugins/no-ai-training-tag.php
    <?php
    /**
     * Plugin Name: No AI training tag
     * Description: Adds noarchive to the robots tag of every page WordPress builds.
     */
    
    add_filter(
    	'wp_robots',
    	function ( $robots ) {
    		$robots['noarchive'] = true;
    
    		return $robots;
    	}
    );
  6. Step 6: Read the tag back

    On a WordPress with no SEO plugin it prints <meta name='robots' content='max-image-preview:large, noarchive' />.

    bash
    curl -sL https://example.com/ | grep -i -o -E "<meta[^>]*name=.robots.[^>]*>"

Four things about those files.

  • No search name is in the group. OAI-SearchBot, Claude-SearchBot, Applebot, meta-webindexer, Amzn-SearchBot, PerplexityBot, Googlebot and Bingbot go on following WordPress's rules for every crawler.
  • Four lines refuse more than training, as said above: meta-externalagent, Amazonbot, CCBot and, for grounding in Gemini Apps, Google-Extended. Delete any line whose cost you do not accept.
  • The list will change. These are the names on the companies' pages on October 8, 2026.
  • WordPress has to be the one answering. The filter runs only where WordPress writes /robots.txt: with permalinks other than Plain, and with no robots.txt file in the site's folder. Where robots.txt is in WordPress covers both. A page cache serves its old copy of a tag until it is cleared.

To undo either, delete its file. For a real file instead, the robots.txt generator writes one from WordPress's own rules. Its training choice names GPTBot, ClaudeBot and Google-Extended, and the other four can be added by hand.

What does not work, or is not documented to

  • A noai meta tag. One page cited here mentions it: Meta's, which says it relies on robots.txt rather than "non-standard formats like NoAI tags". None says its crawler reads one. Google, Bing, Apple and Amazon each list the robots rules they support, and noai is in none of the lists.
  • A line in llms.txt. None of the crawler pages offers that file as a way to refuse training. llms.txt on WordPress covers what the file is.
  • A notice in the footer or the terms. None of the pages says its crawler reads one. Whether a notice carries legal weight is a different question, which this page does not answer.
  • Blocking a company's addresses in place of a robots.txt rule. Anthropic's page says that may not work as an opt-out, because it stops its bot from reading robots.txt.
  • A Content-Usage rule. An IETF working group is writing a way to state a preference such as train-ai=n in robots.txt and in an HTTP header. Both of its documents are Internet-Drafts, the latest versions dated August 19 and September 14, 2026. A draft is a work in progress, not a standard, and none of the company pages mentions the rule. Relying on it today is guesswork.

What a request cannot do

RFC 9309, the standard for robots.txt, says its rules are not a form of access authorization. A tag is the same kind of thing. Both work on a crawler that chooses to read them.

A fetch a person asked for is treated differently. OpenAI, Perplexity, Meta and Amazon each say their user-triggered fetcher may not follow robots.txt. Perplexity and Amazon say theirs is not used for training.

A request is not read at once. OpenAI and Amazon each say about 24 hours, Meta and Perplexity up to 24 hours, and Amazon adds that its crawlers may use a copy of robots.txt cached within the last 30 days.

What turns a crawler away is a block at the server or a CDN. Cloudflare's documentation, last updated on July 1, 2026, describes settings that block AI crawlers by kind, and the guide linked at the top covers them.

Content already collected

None of the pages documents removing content from a model that was already trained on it. What they say is about the future.

  • Anthropic speaks of the site's future materials. A separate help page, dated September 4, 2026, says a site owner can ask for an address to be blocked from Claude's answers that use web search. That is about answers, not training data.
  • Google speaks of future generations of Gemini models.
  • Microsoft's 2023 post says its rule for noarchive applies going forward.
  • Common Crawl publishes a ledger of the legal opt-out requests it has received, so that users of its data know what owners asked to be excluded from its crawls. Its post of September 17, 2025 does not say past archives are changed.
  • OpenAI, Apple, Meta and Amazon: their crawler pages say nothing about content already collected.

Treat an opt-out as counting from the day it is read.

How to check

The curl command above is the first check. The AI crawler access check reads a site's robots.txt for each crawler on its list by name. Of the training names, that list has GPTBot, ClaudeBot and Google-Extended today. For the other four, read the file yourself.

That tool reads robots.txt and nothing else. It cannot see a firewall, a CDN or a host that turns a crawler away, or whether a crawler obeys. The server's access log is the only record of what a crawler was given.

When to get help

If robots.txt as served has lines you cannot trace, or you cannot tell what is refusing a crawler you wanted, hand it over. WP Ministry's AI search optimization is a one-time check of each place an AI crawler can be turned away, from robots.txt to whatever stands in front of the site, with a written list of what to change.

Common questions

Can I refuse AI training and still appear in ChatGPT's answers?

By OpenAI's account, yes. Its page says GPTBot is for training, OAI-SearchBot decides whether a site is shown in ChatGPT's search answers, and each setting is independent. The file above names only the first.

Does blocking Google-Extended take my site out of AI Overviews?

Google says the name does not affect a site's inclusion in Google Search, and AI Overviews are part of Search. It governs Gemini training, and grounding in Gemini Apps and Vertex AI.

Will opting out remove my content from a model that was already trained?

None of the companies' pages says so. Anthropic, Google and Microsoft word their opt-out as applying to future use, and the others say nothing about content already collected.

Is there a meta tag that stops AI training?

For Microsoft and Amazon, noarchive is documented as one, and at Microsoft it also takes the page out of Copilot's answers. A noai tag is not documented as read by any company whose page is cited here.

More on this subject

AI search check, done for you

AI search check is $199. Every place an AI crawler can be blocked, Google's and Bing's AI reports read, and a written list of what to change. It starts with a free diagnosis.