Skip to content

"Blocked by robots.txt" in Search Console: when it is fine, and how to find the rule on WordPress

"Blocked by robots.txt" means Google did not fetch the address because your robots.txt told it not to. For /wp-admin/ and a store's add-to-cart addresses that is intended. If the address is a page you want found, a rule is catching more than it was written for: find the line and change it.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • The status says Google obeyed your robots.txt. Whether that is a problem depends on which addresses are in the list.
  • WordPress keeps crawlers out of /wp-admin/ by itself, so an address or two under this status is normal.
  • A rule matches every address that starts with it. The longest rule that matches decides, and Allow wins a tie.
  • Disallow: /blog also blocks /blog-news/. A trailing slash keeps a rule to one folder.
  • The "Discourage search engines" box has not written "Disallow: /" into robots.txt since WordPress 5.3.
  • A block does not take a page out of Google, and it stops a noindex on that page from being read.

"Blocked by robots.txt" in Search Console's Page indexing report means Google did not fetch the address because your site's robots.txt told it not to. Google's help for the report defines the status in one sentence: "This page was blocked by your site's robots.txt file." Nothing failed. Google asked your site for its rules and followed them.

So the first question is whether you wanted that. For the dashboard under /wp-admin/, and on a store for the addresses that put a product in a cart, you did: WordPress and WooCommerce write those rules themselves. If the address is a page you want people to find, a rule is catching more than it was written for. This page shows how to tell the two apart, how to find the line that decides, and how to correct the usual mistakes on a WordPress site.

What the status says, and what it does not

Search Console prints the reason as "Blocked by robots.txt". Google's help page heads the same entry "URL blocked by robots.txt". They are one status.

  • Its source is "Website". The help page says that, in general, only reasons with that source are yours to fix. Here it means the rule is on your site, not that the rule is wrong. The same page says of this status that addresses blocked by robots.txt are probably blocked on purpose, and to use your judgment.
  • The list is a sample. The help page calls it an example list, limited to 1,000 rows, that does not necessarily show every address.
  • It is not proof the page is out of Google. The help entry adds that a blocked page can, rarely, still be indexed from what Google learns about it elsewhere. The section on the other status, below, covers that.

Is anything wrong? Read the addresses

Open the status and read the list. What each address is tells you whether to act.

The addressWhat to do
/wp-admin/ or anything under itNothing. It is WordPress's own rule.
Contains add-to-cart=, on a storeNothing. They are WooCommerce's own rules.
A sorted, filtered or search address you chose to blockNothing. Filter addresses on a store covers that choice.
A post, page, product or category you want foundFind the line that blocks it, below.
Every address on the siteGo to the first fix.

A WordPress site with nothing added blocks one path. Its robots.txt keeps crawlers out of /wp-admin/ and lets them back in to one file there, admin-ajax.php, and says nothing about any other address. On a store, WooCommerce adds lines to the same answer for add-to-cart addresses and three of its own upload folders. So a short list of those addresses is a site working as designed.

Find the line that blocks an address

Read the file as your site serves it, not as an editor shows it. WordPress answers /robots.txt itself when no file exists, and a plugin can change that answer. Where robots.txt is in WordPress explains where it comes from and how to change it. If the address shows your home page or a 404, the site serves no rules today and the report describes what Google read on an earlier visit; the same guide covers that case.

This fetches the file and prints every Allow and Disallow line whose path the address starts with, each with its line number. Put your site in place of https://example.com, and the address from the report, from its first slash, in place of /blog-news/.

bash
curl -s https://example.com/robots.txt | tr -d '\r' | awk -v path="/blog-news/" '
  tolower($1) ~ /^(allow|disallow):$/ && $2 != "" && index(path, $2) == 1 {
    print NR ": " $0
    if (length($2) > n || (length($2) == n && tolower($1) == "allow:")) { n = length($2); longest = $0 }
  }
  END { print (n ? "Longest match: " longest : "No rule matches " path) }'

For /wp-admin/admin-ajax.php on a fresh WordPress it prints this:

text
2: Disallow: /wp-admin/
3: Allow: /wp-admin/admin-ajax.php
Longest match: Allow: /wp-admin/admin-ajax.php

Two rules match, and the longer one decides, so the file is open. That is the rule of the robots.txt standard, RFC 9309, and of Google's documentation:

  • A rule matches every address that starts with its path. It is a prefix, not a folder.
  • The longest matching rule decides, wherever it sits in the group. Between an Allow and a Disallow of the same length, Google uses the less restrictive one.
  • A crawler follows the group that names it, and the group for * only when none does.
  • Two characters are wildcards. Google reads * as any run of characters and a closing $ as the end of the address.

The command is a quick look, not a parser. It reads every group as one and takes * and $ as plain characters. The robots.txt tester applies all four rules: paste the file, give the address, choose a crawler, and it names the line that decides. It works in your browser. The page indexing check does the same for a live address and reads the page's own tags as well.

In Search Console, inspect the address with the URL Inspection tool. "Crawl allowed?" says whether your site let Google crawl the page or blocked it with a robots.txt rule, and "Test live URL" asks the same of the site as it is now. Google's help entry for the status names a robots.txt tester of its own; its help for the robots.txt report gives the URL Inspection tool as the way to test one address.

What does not cause it: the "Discourage search engines" box

The box labeled "Discourage search engines from indexing this site" is a natural suspect, and it is not the cause. WordPress's reference for do_robots() records that version 5.3 removed the Disallow: / the box used to write. With the box ticked today, /robots.txt still blocks only /wp-admin/. The box puts a noindex tag on every page instead, and Google's help page has a separate reason for a page marked that way. What the box does has the details.

The other status: "Indexed, though blocked by robots.txt"

Google's help page lists this one separately, as a warning and not as a reason for leaving a page out. The rule on your site is the same. The difference is on Google's side: other pages link to the address, and Google indexed it from those links without fetching it. The help page says any snippet for such a result will probably be very limited.

What to do follows from what you want. To keep the page out of Google, take the block away and use a noindex, as in "Unblock a page you want out of Google" above. To have it found, take the block away.

When robots.txt does not explain it

robots.txt is one of several things that can keep a page out of Google. WordPress site not showing up on Google checks the others in order.

If the list holds pages you want found and you cannot trace the rule, a WordPress SEO audit from WP Ministry looks at what robots.txt tells Google's and Bing's crawlers, and whether a rule is blocking something you want found.

What causes it

How to fix it

Remove a robots.txt file left over from a staging copy

  • Easy
  • Low risk
  • About 5 minutes
  • Steps tested on WordPress 7.1.3

A staging copy can be closed to crawlers with a two-line file. Copied to the live site with everything else, it closes that too:

robots.txt
User-agent: *
Disallow: /

Every address starts with /, so the check above prints Longest match: Disallow: / for any page. The site still loads for visitors, because robots.txt is a request to crawlers and blocks nobody.

  1. Step 1: Look in the site's root

    The file sits in the site's root, beside wp-config.php on most installs. If there is no robots.txt there, the line comes from code: see "Find the code that adds the rule" below.

  2. Step 2: Take the file away

    Download a copy, then delete it. Over SSH, from the same folder, this renames it instead:

    bash
    mv robots.txt robots.txt.old
  3. Step 3: Load /robots.txt again

    WordPress answers the address itself once the file is gone. Its answer blocks /wp-admin/ and nothing else, and the check prints No rule matches for your pages.

If the file held rules you want to keep, edit it instead: delete the Disallow: / line and leave the rest.

To undo it: Rename the file back to robots.txt.

Close a rule that is broader than it was meant to be

  • Easy
  • Low risk
  • About 5 minutes
  • Steps tested on WordPress 7.1.3

This file was meant to keep crawlers out of a folder named /blog/:

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /blog

The last rule matches every address that starts with /blog. That is /blog/ and everything in it, and also /blog-news/ and /blogging-tips/, which are pages of their own. Google's documentation gives the same example with /fish, which matches /fishheads. For /blog-news/ the check prints:

text
4: Disallow: /blog
Longest match: Disallow: /blog

A closing slash keeps the rule to the folder:

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /blog/

Now /blog/ and what is under it are blocked, and the check prints No rule matches for the other two. Look for the same mistake wherever a rule ends in a word: /shop also matches /shopping-guide/, and /tag matches /tagline/.

To undo it: Put the old line back.

Stop blocking /wp-content/ and /wp-includes/

  • Easy
  • Low risk
  • About 5 minutes
  • Steps tested on WordPress 7.1.3

A file with these lines is meant to keep crawlers out of WordPress's own folders:

robots.txt
User-agent: *
Disallow: /wp-admin/
Disallow: /wp-includes/
Disallow: /wp-content/

Those folders are not private. This lists the stylesheets and scripts your home page loads from them:

bash
curl -s https://example.com/ | grep -o -E "/wp-(content|includes)/[^\"' ?]+\.(css|js)" | sort -u

A fresh WordPress prints a list of block stylesheets under /wp-includes/ and the theme's own under /wp-content/themes/. The rules above match every one of them. Google's documentation says that when Googlebot finds an address disallowed it skips the request, and that Google Search does not render JavaScript from blocked files. Its introduction to robots.txt says not to block resource files a page depends on, because Google will not do a good job of analyzing pages that need them.

/wp-content/ also holds /wp-content/uploads/, where your images and PDFs are. Google's introduction treats a PDF as a web page, so a blocked PDF is a blocked page, and says a block on image files keeps them out of its results.

Delete the two lines, and keep the exception WordPress writes for admin-ajax.php:

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

That is WordPress's own answer, so deleting the file does the same. To see a page as Google renders it, run "Test live URL" in the URL Inspection tool and open "View tested page". The help page says it shows a screenshot and the resources that were loaded.

To undo it: Put the two lines back.

Find the code that adds the rule

  • Takes care
  • Back up first
  • About 15 minutes
  • Steps tested on WordPress 7.1.3

If /robots.txt shows a rule and no robots.txt file exists, code wrote it. WordPress passes its answer through a filter named robots_txt, and anything attached to it can add lines. This is what such code looks like, here as a must-use plugin, a file WordPress loads on every request:

wp-content/mu-plugins/members-area.php
<?php
// Keeps crawlers out of the members' area.
add_filter( 'robots_txt', function ( $output ) {
	return $output . "\nUser-agent: *\nDisallow: /members/\n";
} );
  1. Step 1: Check an SEO plugin's settings first

    If an SEO plugin has a robots.txt editor, look there before anywhere else, and change the rule on that screen.

  2. Step 2: Search the code for the filter's name

    Over SSH, from the site's root, this prints every file under wp-content that mentions it:

    bash
    grep -rl "robots_txt" wp-content
  3. Step 3: Change the rule where it is written

    A path under wp-content/plugins/ names the plugin: change the rule in its settings, or ask its author, and leave its files alone. A path under wp-content/mu-plugins/ or in the theme's functions.php is your site's own code: take a backup, then delete or correct the line that writes the rule.

  4. Step 4: Load /robots.txt again

    The rule is gone from the answer as soon as the code is.

Must-use plugins do not show in the ordinary plugin list. WordPress's documentation says hosts commonly add them, so the file may not be one you wrote.

To undo it: Restore the file you changed, or the plugin's old setting.

Unblock a page you want out of Google, so its noindex can be read

  • Easy
  • Low risk
  • About 10 minutes
  • Steps tested on WordPress 7.1.3

Blocking a page in robots.txt to remove it from Google is the opposite mistake. Google's introduction to robots.txt says the file is not a way to keep a page out of Google, and the help page for this report says a robots.txt rule will prevent a noindex from being seen. What removes a page is a noindex that Google is allowed to fetch.

WordPress's search results show both signals at once. WordPress has put noindex on them since version 5.7. Suppose robots.txt also carries Disallow: /?s=. For /?s=test the check then prints:

text
4: Disallow: /?s=
Longest match: Disallow: /?s=

And this prints the tag on the same address, which only a request for the page can show:

bash
curl -s "https://example.com/?s=test" | grep -i -o "<meta[^>]*robots[^>]*>"
text
<meta name='robots' content='noindex, follow, max-image-preview:large' />

A crawler that obeys the rule never makes that request, so it never sees the tag. The URL Inspection tool's help says the same from Google's side: for a page blocked by robots.txt, "Indexing allowed?" always reads "Yes", because Google cannot see a noindex there.

Delete the Disallow line for the page, and leave the noindex where it is. The check then prints No rule matches, and the tag is there for Google to read on its next visit. WordPress search pages in Google covers the one case for blocking them anyway.

Never block the cart, checkout or account pages of a store for the same reason: WooCommerce marks them noindex.

To undo it: Put the Disallow line back.

After a change, let Google read the new file

  • Easy
  • No risk
  • About 10 minutes

Google's documentation says it generally caches a robots.txt file for up to 24 hours, so a change is not acted on at once.

  1. Step 1: Check what Google holds

    The robots.txt report, among the property's settings in Search Console, shows the file Google last fetched and when. The help page says you can request a recrawl of the file after a critical change, from the menu beside it, that you generally do not need to, and that a recrawl of the file does not guarantee an immediate recrawl of the addresses it unblocked.

  2. Step 2: Test one address

    In the URL Inspection tool, run "Test live URL" on an address you unblocked. "Crawl allowed?" should no longer say "No".

  3. Step 3: Decide whether to press "Validate fix"

    The help page says validation asks Google to confirm a fix, typically takes up to about two weeks and can take much longer, and stops when Google finds a single remaining instance. An address you blocked on purpose is still an instance, so validation suits a list in which every address was a mistake. Otherwise wait: the same page says Google updates the count whenever it crawls the addresses, validated or not.

An unblocked address is one Google may fetch. Whether it is then indexed is Google's decision, and the address can move to another reason in the report, such as Crawled - currently not indexed.

When to get help

When the list runs to hundreds of addresses, or holds pages you want found and you cannot trace the rule to a file, a plugin's setting or a line of code. Reading the file as it is served, the code that writes it and the report side by side is a short job for someone who does it often.

Common questions

Is "Blocked by robots.txt" an error?

No. It reports that Google followed a rule on your site. It is a problem only when the address is one you want in Google.

Why is /wp-admin/ in the list?

Because WordPress's own robots.txt asks crawlers to stay out of the dashboard, and Google came across the address. Leave the rule as it is.

I changed the rule. Why is the address still listed?

Google works from a stored copy of robots.txt for a while, and the report's help says the count changes when Google crawls the address again. Run "Test live URL" on the address to see what the site answers now.

My /robots.txt answers 404. Is that blocking Google?

No. Google's documentation says it treats a 4xx answer for robots.txt, other than 429, as if there were no file and no restrictions.

I only added rules for AI crawlers. Can they block Google?

Not if each rule sits in a group that names the AI crawler. Googlebot follows a group that names it and otherwise the group for *, so a rule added under User-agent: * applies to Google too. Blocking or allowing AI crawlers on WordPress has groups written for each crawler.

More on this subject

Would you rather we fixed it?

SEO audit is $249. A written audit, each finding with its source, and up to 2 hours of fixes done and tested again. It starts with a free diagnosis.