How to stop AI training on your WordPress content and stay in AI search answers
You can ask the crawlers that collect for AI training to stay out, one company at a time, and stay in most assistants' search answers. The request is a robots.txt group or, at Microsoft and Amazon, a noarchive tag. It is not a lock, and it does not take back what was already collected.
- By
- WP Ministry
- Published
- Tested on
- WordPress 7.1.3, PHP 8.3.35
In short
- OpenAI, Anthropic, Apple, Amazon and Meta give training and search different names in robots.txt, so one can be refused and the other left alone.
- Google-Extended governs Gemini training and does not affect Google Search. AI Overviews are part of Search and follow Search's own controls.
- Microsoft has no training name in robots.txt. Its opt-out is a noarchive tag on the page, which also takes the page out of Copilot's answers.
- Every opt-out is a request about future use. No page cited here says content already collected is taken back out.
- A noai tag, a line in llms.txt and a notice in the footer are not documented as read by any of these crawlers.
- The list of names changes. Read each company's page again before relying on any list, this one included.
You can ask the crawlers that collect pages for AI training to stay out of a WordPress site, one company at a time, and each company below that offers an opt-out says it honors it. OpenAI, Anthropic, Apple, Amazon and Meta use one name for training and another for the search behind their assistant, so a site can refuse the first and stay in the second.
Three limits come with that. A request is not a lock. It is about future use: none of the pages cited here says content already collected is taken back out. And at Microsoft the two are tied together, because the tag that refuses training also takes a page out of Copilot's answers.
How to block or allow AI crawlers on WordPress covers how a robots.txt group works and where a CDN, a host or a plugin can block a crawler. This page stays on one question: training out, search answers in.
What each company documents
Each row is from the company's own page, read on October 8, 2026. That is the standing of everything in the table: documented by the company, about its own crawler. Whether a crawler does what its page says is not shown here, and no study of that is cited. A name in the table is one to write after User-agent: in robots.txt, unless the cell says it is a tag.
| Company | What governs training | What governs its search answers | Separate? | What the company says a block does | Its page |
|---|---|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | Yes. It says each setting is independent | Disallowing GPTBot indicates the content should not be used to train its models. Without OAI-SearchBot a site is not shown in ChatGPT's search answers, though it can still appear as a navigational link | Undated |
| Anthropic | ClaudeBot | Claude-SearchBot | Yes | Restricting ClaudeBot signals that the site's future materials should be left out of its training data. Disabling Claude-SearchBot may reduce the site's visibility and accuracy in users' search results | April 7, 2026 |
Google-Extended | Googlebot, and Search's own controls | Partly. See below | Google-Extended governs training of future Gemini models, and grounding in Gemini Apps and Vertex AI. It does not affect inclusion or ranking in Search | July 14, 2026 | |
| Microsoft | A noarchive or nocache robots tag on the page. Its pages give no robots.txt name for training | Bingbot, and the same two tags | No. One tag governs both | noarchive: not used for training, and not linked in Copilot. nocache: only the address, title and snippet are used, for both | Help page undated |
| Apple | Applebot-Extended | Applebot. For AI answers, a nosnippet tag | Yes | Disallowing Applebot-Extended opts out of training Apple's foundation models. The pages can still be in search results, and the rule is not considered in ranking | September 4, 2026 |
| Meta | meta-externalagent | meta-webindexer | Two names, but the first is described as for training models or indexing content directly | Allowing meta-webindexer helps Meta cite and link the site in Meta AI's responses. The page does not say what refusing meta-externalagent costs | May 21, 2026 |
| Amazon | Amazonbot, or a noarchive tag on the page | Amzn-SearchBot | Yes. It says each setting is independent | Amazonbot improves its products and may be used to train its models. noarchive means the page is not used for model training. Amzn-SearchBot makes content eligible for search experiences such as Alexa and does not crawl for training | Undated |
| Perplexity | Nothing. It says neither of its crawlers collects for training foundation models | PerplexityBot | Nothing to separate | It recommends allowing PerplexityBot for a site to appear in its search results | Undated |
| Common Crawl | CCBot | It has no assistant. Others build on its open archive | No. One crawler, one archive | Its crawler stops crawling the site. Its post of February 13, 2024 says that also keeps the site from those who build search indexes on the data | Undated |
Two things in the table are easy to miss.
Two of the names are not crawlers. Google-Extended and Applebot-Extended send no requests. Google and Apple each say the name only decides how pages their ordinary crawler fetched may be used, so neither shows in an access log.
Three of the names do more than training. meta-externalagent, Amazonbot and CCBot each serve other uses too, as their rows say. Refusing one refuses those uses with it.
Google: three controls for three things
- Training Gemini. Google's page on its crawlers says
Google-Extendedgoverns whether content Google crawls may be used to train future generations of Gemini models and to ground answers in Gemini Apps and Vertex AI. It says the name does not affect a site's inclusion in Google Search and is not a ranking signal there. No name covers training alone. - AI Overviews and AI Mode. These are part of Search. Google's page on AI features says AI is built into Search, so the rules for
Googlebotare the control for how a site is crawled for it. No crawler name admits Search and refuses AI Overviews. - How much of a page Search may show or use. Google's specification says
nosnippetprevents a page's content from being used as a direct input for AI Overviews and AI Mode, andmax-snippetlimits how much may be. The cost is in the same sentence: both apply to every kind of result, so the ordinary snippet under the page's title goes or shrinks too.data-nosnippetkeeps one part of a page out of snippets.
Search Console also has a setting that takes a whole site out of AI Overviews and AI Mode. Its help page says the setting does not affect AI training, and names Google-Extended as the way to limit training of the models that generate those answers. How to appear in Google AI Overviews covers that setting and the snippet rules.
So at Google, refusing training is one line in robots.txt, and by Google's account it costs nothing in Search.
Microsoft: a tag on the page, not a name in robots.txt
Bing's help page on robots tags gives the two that matter here.
| Tag | In Copilot | For training Microsoft's models |
|---|---|---|
noarchive | Not linked | Not used |
nocache | Only the address, snippet and title are shown | Only addresses, titles and snippets are used |
| Both | Treated as nocache | Treated as nocache |
Bing's post of September 22, 2023, which introduced this, says a page with either tag still appears in Bing's search results. Bing's Webmaster Guidelines state the cost: noarchive prevents content from being used in Copilot's responses, and nocache reduces how much of a page Copilot can cite.
So at Microsoft the two cannot both be had in full. noarchive is the whole opt-out and gives up Copilot. nocache keeps a short form of both.
Two more companies document the same tag. Amazon's page says its crawlers read noarchive as an instruction not to use the page for model training. Google's specification lists noarchive and nocache among rules Google Search does not use and ignores.
The WordPress side
WordPress writes /robots.txt and each page's robots tag itself, and has a filter for each. A small must-use plugin on a filter adds to what WordPress writes and removes nothing.
Step 1: Create the mu-plugins folder if it is not there
A fresh WordPress has no
mu-pluginsfolder inwp-content, so create it. WordPress loads every PHP file in that folder on each request, with nothing to activate.Step 2: Add the robots.txt group
Save this as
no-ai-training.phpin that folder. It adds one group with the seven training names from the table.wp-content/mu-plugins/no-ai-training.php<?php /** * Plugin Name: No AI training * Description: Asks the crawlers that collect for AI training to stay out of the site. */ add_filter( 'robots_txt', function ( $output, $public ) { $output .= "\nUser-agent: GPTBot\n"; $output .= "User-agent: ClaudeBot\n"; $output .= "User-agent: Google-Extended\n"; $output .= "User-agent: Applebot-Extended\n"; $output .= "User-agent: meta-externalagent\n"; $output .= "User-agent: Amazonbot\n"; $output .= "User-agent: CCBot\n"; $output .= "Disallow: /\n"; return $output; }, 10, 2 );Step 3: Read robots.txt as served
Replace
example.comwith your site's address.bashcurl -s https://example.com/robots.txtStep 4: Check what came back
WordPress's own three lines and its Sitemap line are still there, and the new group follows them.
textUser-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/wp-sitemap.xml User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: meta-externalagent User-agent: Amazonbot User-agent: CCBot Disallow: /Step 5: Add the tag, if you accept its cost
Save this as
no-ai-training-tag.phpin the same folder. It addsnoarchiveto the robots tag of every page WordPress builds. To keep the short form of Copilot's answers, writenocachein place ofnoarchivein the line that sets it.wp-content/mu-plugins/no-ai-training-tag.php<?php /** * Plugin Name: No AI training tag * Description: Adds noarchive to the robots tag of every page WordPress builds. */ add_filter( 'wp_robots', function ( $robots ) { $robots['noarchive'] = true; return $robots; } );Step 6: Read the tag back
On a WordPress with no SEO plugin it prints
<meta name='robots' content='max-image-preview:large, noarchive' />.bashcurl -sL https://example.com/ | grep -i -o -E "<meta[^>]*name=.robots.[^>]*>"
Four things about those files.
- No search name is in the group.
OAI-SearchBot,Claude-SearchBot,Applebot,meta-webindexer,Amzn-SearchBot,PerplexityBot,GooglebotandBingbotgo on following WordPress's rules for every crawler. - Four lines refuse more than training, as said above:
meta-externalagent,Amazonbot,CCBotand, for grounding in Gemini Apps,Google-Extended. Delete any line whose cost you do not accept. - The list will change. These are the names on the companies' pages on October 8, 2026.
- WordPress has to be the one answering. The filter runs only where WordPress writes
/robots.txt: with permalinks other than Plain, and with norobots.txtfile in the site's folder. Where robots.txt is in WordPress covers both. A page cache serves its old copy of a tag until it is cleared.
To undo either, delete its file. For a real file instead, the robots.txt generator writes one from WordPress's own rules. Its training choice names GPTBot, ClaudeBot and Google-Extended, and the other four can be added by hand.
What does not work, or is not documented to
- A
noaimeta tag. One page cited here mentions it: Meta's, which says it relies on robots.txt rather than "non-standard formats like NoAI tags". None says its crawler reads one. Google, Bing, Apple and Amazon each list the robots rules they support, andnoaiis in none of the lists. - A line in llms.txt. None of the crawler pages offers that file as a way to refuse training. llms.txt on WordPress covers what the file is.
- A notice in the footer or the terms. None of the pages says its crawler reads one. Whether a notice carries legal weight is a different question, which this page does not answer.
- Blocking a company's addresses in place of a robots.txt rule. Anthropic's page says that may not work as an opt-out, because it stops its bot from reading robots.txt.
- A
Content-Usagerule. An IETF working group is writing a way to state a preference such astrain-ai=nin robots.txt and in an HTTP header. Both of its documents are Internet-Drafts, the latest versions dated August 19 and September 14, 2026. A draft is a work in progress, not a standard, and none of the company pages mentions the rule. Relying on it today is guesswork.
What a request cannot do
RFC 9309, the standard for robots.txt, says its rules are not a form of access authorization. A tag is the same kind of thing. Both work on a crawler that chooses to read them.
A fetch a person asked for is treated differently. OpenAI, Perplexity, Meta and Amazon each say their user-triggered fetcher may not follow robots.txt. Perplexity and Amazon say theirs is not used for training.
A request is not read at once. OpenAI and Amazon each say about 24 hours, Meta and Perplexity up to 24 hours, and Amazon adds that its crawlers may use a copy of robots.txt cached within the last 30 days.
What turns a crawler away is a block at the server or a CDN. Cloudflare's documentation, last updated on July 1, 2026, describes settings that block AI crawlers by kind, and the guide linked at the top covers them.
Content already collected
None of the pages documents removing content from a model that was already trained on it. What they say is about the future.
- Anthropic speaks of the site's future materials. A separate help page, dated September 4, 2026, says a site owner can ask for an address to be blocked from Claude's answers that use web search. That is about answers, not training data.
- Google speaks of future generations of Gemini models.
- Microsoft's 2023 post says its rule for
noarchiveapplies going forward. - Common Crawl publishes a ledger of the legal opt-out requests it has received, so that users of its data know what owners asked to be excluded from its crawls. Its post of September 17, 2025 does not say past archives are changed.
- OpenAI, Apple, Meta and Amazon: their crawler pages say nothing about content already collected.
Treat an opt-out as counting from the day it is read.
How to check
The curl command above is the first check. The AI crawler access check reads a site's robots.txt for each crawler on its list by name. Of the training names, that list has GPTBot, ClaudeBot and Google-Extended today. For the other four, read the file yourself.
That tool reads robots.txt and nothing else. It cannot see a firewall, a CDN or a host that turns a crawler away, or whether a crawler obeys. The server's access log is the only record of what a crawler was given.
When to get help
If robots.txt as served has lines you cannot trace, or you cannot tell what is refusing a crawler you wanted, hand it over. WP Ministry's AI search optimization is a one-time check of each place an AI crawler can be turned away, from robots.txt to whatever stands in front of the site, with a written list of what to change.
Common questions
Can I refuse AI training and still appear in ChatGPT's answers?
By OpenAI's account, yes. Its page says GPTBot is for training, OAI-SearchBot decides whether a site is shown in ChatGPT's search answers, and each setting is independent. The file above names only the first.
Does blocking Google-Extended take my site out of AI Overviews?
Google says the name does not affect a site's inclusion in Google Search, and AI Overviews are part of Search. It governs Gemini training, and grounding in Gemini Apps and Vertex AI.
Will opting out remove my content from a model that was already trained?
None of the companies' pages says so. Anthropic, Google and Microsoft word their opt-out as applying to future use, and the others say nothing about content already collected.
Is there a meta tag that stops AI training?
For Microsoft and Amazon, noarchive is documented as one, and at Microsoft it also takes the page out of Copilot's answers. A noai tag is not documented as read by any company whose page is cited here.
- GuideBing Webmaster Tools for WordPress: setup, IndexNow, and what it means for ChatGPT and Copilot
- GuideGenerative engine optimization (GEO): what is documented, what was studied and what is guesswork
- GuideHow to appear in Google AI Overviews: what Google requires and what keeps a WordPress page out
- GuideHow to block AI crawlers on WordPress, or let them in: robots.txt, Cloudflare, hosts and plugins
- Guidellms.txt on WordPress: what it is, who reads it and whether your site needs one
- GuideAI bots slowing down your website: how to confirm it on WordPress and what to do, mildest first

