AI bots slowing down your website: how to confirm it on WordPress and what to do, mildest first
Count the requests in the server's access log by user agent. If AI crawlers lead, ask them in robots.txt to stay out of the addresses that cost the most, answer named crawlers with 429 while the server is overloaded, and block one outright only if it keeps coming.
- By
- WP Ministry
- Published
- Tested on
- WordPress 7.1.3, PHP 8.3.35
In short
- The access log settles it. Count requests by user agent, then see what the leading crawler asked for, before changing anything.
- A user agent is a claim. Most of the companies publish IP addresses or a reverse DNS name to check it against.
- Anthropic, Common Crawl and Bing document obeying Crawl-delay. Google, Amazon and Apple say they do not. OpenAI's, Perplexity's and Meta's pages do not mention it.
- A 429 answered by the web server to named crawlers leaves visitors alone and does not start WordPress.
- Refusing a search crawler outright takes the site out of that assistant's search answers, by the companies' own accounts.
- Never slow Googlebot or Bingbot with a 403. Google's documented way is a 500, 503 or 429, for a day or two at most.
When a host warns about CPU, or pages turn slow or answer 503, the access log shows whether GPTBot, ClaudeBot, Bytespider, Amazonbot or another AI crawler is behind it. Count the requests by user agent and look at what the leading crawler asked for. Then act in order of what it costs you. Ask the crawlers in robots.txt to stay out of the addresses that are expensive to build. Answer named crawlers with a 429 status while the server is struggling. Refuse one outright if it keeps coming. Put a cache or a CDN in front so that each request costs less.
This page is about load. Whether to let AI crawlers read the site at all is a separate decision, covered in how to block AI crawlers on WordPress, or let them in, and keeping content out of training is in opting out of AI training.
Confirm it in the access log
The web server writes one line to its access log for every request it answers. The server's configuration sets where the file is, so the path differs from host to host. In cPanel, each domain's log is downloaded from the Raw Access page. Elsewhere, look for an item named for logs, or ask the host.
The commands read a log in the "combined" format, which Apache's documentation defines. Each line starts with the IP address the request came from, has the address that was asked for as its seventh item, and ends with the user agent in quotation marks. Replace /path/to/access.log with the log's path.
Step 1: Count requests by user agent
Each line of the result is a count and a user agent, largest first. A crawler names itself inside the string, as in
GPTBot/1.4. If the top lines are browsers, go to the last section of this page.bashawk -F'"' '{ print $6 }' /path/to/access.log | sort | uniq -c | sort -rn | head -n 20Step 2: See what one crawler asked for
Replace
GPTBotwith the name that leads your list. The result is a count and an address. Look for addresses that differ only after the?, such as/?s=and?orderby=.bashgrep -i 'GPTBot' /path/to/access.log | awk '{ print $7 }' | sort | uniq -c | sort -rn | head -n 20Step 3: Count requests by IP address
To list only the IP addresses one crawler's name came from, change
$7to$1in the second command.bashawk '{ print $1 }' /path/to/access.log | sort | uniq -c | sort -rn | head -n 20
Check that a crawler is who it says
A user agent is text the requester chooses. Common Crawl's page warns of crawlers falsely identifying themselves as CCBot, and Bing's says the string "can be easily spoofed by anyone". The IP address a request came from cannot be borrowed that way. Each row below is from the company's own page, read on October 8, 2026.
| Company and the names in a log | Crawl-delay in robots.txt | How to check a request is theirs |
|---|---|---|
OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User | Not mentioned, nor any other way to ask for a slower crawl | A list of IP addresses published for each crawler |
Anthropic: ClaudeBot, Claude-SearchBot, Claude-User | Supported. Its example is Crawl-delay: 1 | A published list of IP addresses |
Perplexity: PerplexityBot, Perplexity-User | Not mentioned | A published list for each. It tells firewall owners to match the name and the IP address together |
Google: Googlebot | Not supported. Answer 500, 503 or 429 instead (see below) | Reverse DNS to a name ending in googlebot.com for its common crawlers, and published lists |
Microsoft: bingbot | Honored, from 1 to 20 seconds. Crawl Control in Bing Webmaster Tools sets a slower rate by the hour | Reverse DNS to a name ending in search.msn.com, or Bing's Verify Bingbot tool |
Meta: meta-externalagent, meta-webindexer, meta-externalfetcher | Not mentioned | The page gives the user agent strings and no list of IP addresses |
Amazon: Amazonbot, Amzn-SearchBot, Amzn-User | "They do not support the crawl-delay directive." | A published list for each |
Apple: Applebot | "Applebot does not follow crawl-delay." Apple says its crawl rate adjusts when a site slows down or returns errors | Reverse DNS in applebot.apple.com, or a published list |
Common Crawl: CCBot | Obeyed. It also slows down when a server answers 429 or a 5xx status | Reverse DNS to a name ending in crawl.commoncrawl.org, and a published list |
ByteDance: Bytespider | No page of ByteDance's that documents the crawler was found | Nothing published that we could find |
Bytespider is attributed to ByteDance in a report Cloudflare published on July 1, 2025.
A published list is a file of address ranges to compare an IP address against. A reverse DNS check asks which name an IP address belongs to, then asks for that name's address and confirms it is the same one. A request that carries a crawler's name from an IP address that fails the check is not that company's crawler, and nothing the company says applies to it.
Why crawlers cost a WordPress site more than visitors do
WordPress builds a page by running PHP and querying the database. A page cache saves the finished page as a static file, and WordPress's documentation says serving those files reduces the processing load on the server. A cache can only serve an address it has already stored, and two kinds of address are not there when a crawler asks:
- Addresses with no end. WordPress answers
/?s=followed by any word with a search results page. A store adds a sorting and a filter address to every list of products, and a calendar adds one for every date. - Addresses that must not be cached. WooCommerce's documentation says the cart, checkout and account pages have to stay dynamic, because they show one customer's own data.
A visitor opens a few of these. A crawler follows every link it finds. Google's documentation describes the effect on its own crawler: filter addresses "can generate infinite URL spaces", and it lists faceted navigation and calendars among the most common causes of a sharp increase in crawling.
The Wikimedia Foundation, which runs Wikipedia, reported the same pattern on its own servers on April 1, 2025. Crawlers "bulk read" pages that few people visit and that are not in its cache, so bots made at least 65% of its most expensive traffic while making about 35% of its page views. That is one operator's account of its own sites, from an organization that also sells API access to its content. The measurement of your site is the second command above.
WooCommerce filter URLs lists the addresses a store answers at, and WordPress search pages in Google covers /?s=.
First: ask them to stay out of the expensive addresses
The measures are in order, mildest first. This one costs nothing, and works only on a crawler that honors robots.txt.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xml
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Amazonbot
User-agent: meta-externalagent
User-agent: Bytespider
Crawl-delay: 10
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /*?s=
Disallow: /*?*orderby=
Disallow: /*?*filter_
Disallow: /*?*add-to-cart=The first group is what a fresh WordPress answers with. A crawler that a file names follows only its own group, which is why the second group repeats the /wp-admin/ lines. A file in the site's root replaces WordPress's answer, so if yours comes from WordPress or a plugin, add the second group the way WordPress robots.txt describes. The robots.txt tester shows which rule decides for an address and a crawler.
Three limits:
Crawl-delayis not in the standard. RFC 9309 definesUser-agent,AllowandDisallow, and lets a crawler interpret other lines if it chooses. Of the six names above, only Anthropic's and Common Crawl's pages say theirs obeys it. Neither says what the number counts. Bing's page, for Bingbot, defines it as the seconds in which one address is crawled.- It is not read at once. OpenAI's and Amazon's pages say about 24 hours, and Perplexity's and Meta's up to 24.
- Googlebot and Bingbot are left out on purpose. WordPress marks its search pages
noindex, and Google's documentation says a crawler blocked by robots.txt never sees that tag.
Run the second log command again after a day. A crawler still asking for the disallowed addresses is not honoring the file.
Second: answer named crawlers with 429 while the server is overloaded
Status 429 means too many requests in a given amount of time. Google's documentation says its crawlers treat it as a sign that the server is overloaded and slow down. Common Crawl says CCBot slows down on a 429 or a 5xx status, and Apple says Applebot adjusts when a site returns errors. The pages of OpenAI, Anthropic, Perplexity, Meta and Amazon do not say what their crawlers do on receiving one. Whatever the crawler does next, the request it just made has cost the server very little.
Step 1: Add the rule above WordPress's own
On Apache, add these lines to the
.htaccessfile in the site's root, above the line# BEGIN WordPress. WordPress rewrites whatever sits between its two markers.The first condition matches a name anywhere in the user agent, in any letter case. Apache's documentation says a status outside the 300s given to the
Rflag is returned as it is, and nothing else is served.GPTBotdoes not matchOAI-SearchBot, andClaudeBotdoes not matchClaude-SearchBot, so the search crawlers are still served. A shorter word such asClaudewould match all three of Anthropic's.The second condition leaves
/robots.txtalone. RFC 9309 says a crawler that gets a 4xx status for robots.txt may treat the site as having no rules at all. It reads the request as it was sent, and not%{REQUEST_URI}, because WordPress answers/robots.txtthroughindex.php: written the other way, the rule answers 429 for robots.txt too..htaccess# Ask named crawlers to come back later. Remove these lines when the load has passed. <IfModule mod_rewrite.c> RewriteEngine On RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|Amazonbot|meta-externalagent|Bytespider) [NC] RewriteCond %{THE_REQUEST} !\s/robots\.txt[\s?] RewriteRule ^ - [R=429,L] </IfModule>Step 2: Ask for a page under a crawler's name
Replace
example.comwith your site's address. The command sends OpenAI's published example string forGPTBotand prints the status that comes back.It prints
429. The answer is Apache's own short error page, and WordPress is not started. Withrobots.txtadded to the end of the address, it prints200.bashcurl -s -o /dev/null -w '%{http_code}\n' -A 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot' https://example.com/Step 3: Ask again as a browser
It prints
200. Visitors are not affected.bashcurl -s -o /dev/null -w '%{http_code}\n' -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Safari/605.1.15' https://example.com/
To answer 503, which means the server cannot handle the request for now, change 429 to 503 in the rule.
Either status may carry a Retry-After header, and Google's page says it tells its crawlers when they can retry. The rule above does not send it. On Apache the header is set with a Header directive, which needs the mod_headers module. Ask the host whether it is loaded before adding one: a directive for a module that is not loaded makes every page answer 500.
nginx has no file like .htaccess. WordPress's documentation says its configuration is done at the server level by an administrator, so there the rule is the host's to add.
Third: refuse a crawler outright
For a crawler that ignores the first two measures, change the answer to a refusal. Put this in the same place in .htaccess, with only the names you mean to turn away.
# Refuse named crawlers.
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (GPTBot|Bytespider) [NC]
RewriteCond %{THE_REQUEST} !\s/robots\.txt[\s?]
RewriteRule ^ - [F]
</IfModule>Apache's documentation says the [F] flag returns 403 Forbidden. The two commands above then print 403 and 200.
A rule that goes by name stops only requests that carry the name.
Fourth: stop them before they reach the server
A rule in .htaccess still makes the server answer every request. A CDN or a host's firewall answers in the server's place. Cloudflare's documentation describes two tools for it. AI Crawl Control lists each AI crawler with its request counts and lets you allow or block each one. Rate limiting rules cap how many requests matching an expression are let through in a period, and the same page warns that applying one to verified bots might affect search. How to set up a CDN covers putting one in front of a site, and WordPress firewalls compared covers the choices. Some hosts filter AI crawlers on their own servers, as the page on blocking or allowing them describes, so ask yours.
Fifth: make each request cost less
A page cache turns the crawl of ordinary pages into the serving of files. WordPress caching plugins compared covers choosing one, and server response time shows how to see whether it is working. A cache does nothing for an address it has not stored, which is why the first measure comes first.
What not to do
- Do not refuse Googlebot or Bingbot. Google's documentation says not to use 401 or 403 to limit its crawl rate. To slow Googlebot urgently, it says to answer 500, 503 or 429 for a couple of hours or a day or two, and warns that addresses answering that way for several days may be dropped from its index. Even a short slowdown means fewer new pages discovered and existing ones refreshed less often. For Bingbot, Bing offers Crawl Control and
crawl-delay. - Do not count on robots.txt for a fetch a person asked for, or for a crawler that promises nothing. OpenAI, Perplexity, Meta and Amazon each say the fetcher they send at a user's request may not follow robots.txt. ByteDance publishes nothing we could find about Bytespider. The standard itself says its rules "are not a form of access authorization".
- Do not block IP ranges copied from an article. Perplexity's page says its lists are updated regularly and to use its own endpoints. Anthropic says blocking its IP addresses may not work as an opt-out, because it stops the crawler from reading robots.txt.
When it is not a crawler
The same commands show when the first guess is wrong.
- One or two IP addresses make most of the requests, under a browser's user agent. That is a program that does not name itself. If it is asking for
wp-login.phporxmlrpc.php, see brute force attacks. - A crawler's name comes from IP addresses that fail the company's check. It is something else using the name. A rule by name still stops it.
- Most requests are for
admin-ajax.php. See admin-ajax.php high CPU usage. - The log is quiet and the CPU is not. Requests are not the cause. Start with how to speed up a slow WordPress site, or whether the fault is the host or the site.
When to get help
Ask the host for the access log if you cannot reach it, and whether it filters crawlers itself. WP Ministry, which publishes this page, sells a WordPress maintenance service: a monthly care plan that keeps one WordPress site updated, backed up, monitored and scanned. On its Care plan and above, a fix like the rules on this page comes out of the month's allowance of time for fixes and small edits.
Common questions
Does Crawl-delay slow down AI bots?
Only the ones that say they obey it. Anthropic's page says its bots support it, and Common Crawl's says CCBot obeys it. Amazon's and Apple's pages say their crawlers do not, Google's says it is not supported, and OpenAI's, Perplexity's and Meta's pages do not mention it. The robots.txt standard does not define it.
Should I answer 429 or 503?
Either. Google's documentation treats them alike for slowing its crawler, and Common Crawl's FAQ names 429 and 5xx together.
Does any of this affect Google Search?
Not if Googlebot is left out of the rules, as it is in every example here. None of the six names in them is a Google crawler.
How do I know a rule is working?
Ask for a page under the crawler's name, as in the second step, and read the access log the next day. The status is the ninth item on each line in the combined format, so a line for a slowed crawler shows 429 there.
- GuideGenerative engine optimization (GEO): what is documented, what was studied and what is guesswork
- GuideHow to appear in Google AI Overviews: what Google requires and what keeps a WordPress page out
- GuideHow to block AI crawlers on WordPress, or let them in: robots.txt, Cloudflare, hosts and plugins
- GuideHow to stop AI training on your WordPress content and stay in AI search answers
- Guidellms.txt on WordPress: what it is, who reads it and whether your site needs one
- GuideBing Webmaster Tools for WordPress: setup, IndexNow, and what it means for ChatGPT and Copilot

