Skip to content

WordPress robots.txt: where it is, what it says by default and how to edit it

A fresh WordPress has no robots.txt file. WordPress answers the address itself, with three lines about /wp-admin/ and a Sitemap line. Change it with a real file in the site's root, which replaces that answer, or with the robots_txt filter, which adds to it.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • A fresh WordPress has no robots.txt file to find. WordPress writes the answer each time the address is requested.
  • The default is three lines about /wp-admin/, and a Sitemap line while the site is open to search engines.
  • A real robots.txt file in the site's root replaces WordPress's answer entirely, the Sitemap line included.
  • The robots_txt filter keeps WordPress's lines and adds yours. It does nothing while a real file exists.
  • With Plain permalinks WordPress does not answer /robots.txt at all.
  • robots.txt controls crawling, not indexing. Google can still list a blocked address, and never reads a noindex behind a block.

A fresh WordPress has no robots.txt file. When a crawler asks for /robots.txt and no file is there, WordPress writes the answer itself: three lines that keep crawlers out of the dashboard, and a line that says where the sitemap is. That is why the address answers and the folder holds nothing by that name.

There are two real ways to change it. A file named robots.txt in the site's root replaces WordPress's answer entirely. A few lines of PHP on the robots_txt filter add to WordPress's answer and keep the rest. An SEO plugin's editor is one of those two underneath.

Why there is no robots.txt file

WordPress treats /robots.txt as an address it serves, like a post or a feed. A rewrite rule sends the request to WordPress, and a function called do_robots() prints the lines. Nothing is saved to disk.

Three things have to be true for WordPress to answer:

  • No real file is in the way. The web server looks for a file first. If robots.txt exists in the site's root, the server sends it and WordPress is never asked.
  • Permalinks are not set to Plain. The rule that catches /robots.txt is one of WordPress's rewrite rules, and with Plain permalinks there are none. The section after the next one covers that case.
  • WordPress is at the root of the address. WordPress's source adds the rule only when the site's home has no folder in it, so a WordPress at example.com/blog/ writes no robots.txt answer. Crawlers would not look in that folder anyway: Google's documentation says they do not check for robots.txt files in subdirectories.

What a fresh WordPress answers with

This is the whole answer on a new site with nothing added, with example.com standing in for the site's address:

text
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml
LineWhat it means
User-agent: *The rules below are for every crawler that no other group names. This is the only group WordPress writes.
Disallow: /wp-admin/Do not crawl anything whose path starts with /wp-admin/, which is the dashboard.
Allow: /wp-admin/admin-ajax.phpOne exception inside that folder. Google applies the rule with the longer path when two match, so this file stays open.
Sitemap: https://example.com/wp-sitemap.xmlWhere the sitemap index is. Google's documentation says a sitemap line belongs to no group and may be read by every crawler.

The exception is there because admin-ajax.php is not only for the dashboard. WordPress's plugin handbook sends every Ajax request to that file, from public pages too, and the change that added the line to WordPress gives that as its reason.

The Sitemap line is not part of do_robots(). It is added by WordPress's own sitemap feature, which arrived in WordPress 5.5, and only in two cases together: the sitemap is switched on, and the site is open to search engines. Anything that switches WordPress's sitemap off takes the line with it. The sitemap at wp-sitemap.xml covers the sitemap itself.

Ticking "Discourage search engines from indexing this site" under Settings, then Reading, removes the Sitemap line and changes nothing else here: since WordPress 5.3 that box no longer writes Disallow: / into this answer, and what the box does instead has its own page.

If the address shows your home page or a 404

On a site with Plain permalinks, WordPress does not answer /robots.txt. A new WordPress can start that way, so check before you look for anything else. In the dashboard it is under Settings, then Permalinks: "Plain" is the first choice in the list. Over SSH, WP-CLI prints the structure in use, and prints an empty line when it is Plain.

bash
wp option get permalink_structure

What a crawler receives then depends on the server. On an Apache server whose .htaccess still holds WordPress's rewrite rules, the request reaches WordPress, which redirects /robots.txt to /robots.txt/ and answers with the home page. With no such rules, the server answers with its own 404. In both cases WordPress's rules are still there at /?robots=1, an address no crawler asks for.

Neither outcome blocks anything. Google's documentation says a 404 for robots.txt is treated as no restrictions at all, and that when HTML comes back in place of rules, the lines that are not rules are ignored. What is lost is the Sitemap line and any rule you meant to add with the filter.

To get WordPress's answer back, choose any structure other than Plain. That changes the address of every post on a site that already has visitors, so read how to change permalinks safely first. A real robots.txt file is served whatever the permalinks are.

Three ways to change robots.txt

WayWhere it livesWordPress's own linesThe Sitemap line
A real filerobots.txt in the site's rootGone. The file is the whole answer.Only if you write it in
The robots_txt filterA small plugin, or the theme's functions.phpKept. Your code adds to them or changes them.Kept, and written by WordPress
An SEO plugin's editorThe plugin's screensDepends on which of the two the plugin usesDepends on the plugin

A real file in the site's root

Create a plain text file named robots.txt, in lowercase, and upload it to the folder that holds wp-admin, wp-content and wp-includes. Google's documentation asks for UTF-8 text made in a text editor, not a word processor, and for the file to sit at the root of the host it applies to.

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /example-private/

Sitemap: https://example.com/wp-sitemap.xml
Replace example.com with your site's address and /example-private/ with the path you want crawlers to stay out of.

From the moment the file is there, it is the entire answer. WordPress's three lines and its Sitemap line are not merged in, which is why the sample repeats them. Leave them out and they are gone. The Sitemap line has to be the full address of the sitemap your site really serves, and from now on it is yours to keep correct.

Delete the file and WordPress answers again, as before.

The robots_txt filter

A filter lets you change the answer WordPress writes, so its own lines stay and the Sitemap line keeps following the site's settings. The smallest place to put one is a must-use plugin: a single PHP file in wp-content/mu-plugins/ that WordPress loads on every request, with nothing to activate.

  1. Step 1: Create the mu-plugins folder if it is not there

    Look in wp-content. If there is no folder named mu-plugins, create one. A fresh WordPress does not have it.

  2. Step 2: Add the file

    Save this as robots-txt-rules.php in that folder. It adds a group of its own after everything WordPress wrote.

    wp-content/mu-plugins/robots-txt-rules.php
    <?php
    /**
     * Plugin Name: robots.txt rules
     * Description: Adds this site's own rules to the robots.txt that WordPress answers with.
     */
    
    add_filter(
    	'robots_txt',
    	function ( $output, $public ) {
    		$output .= "\nUser-agent: *\n";
    		$output .= "Disallow: /example-folder/\n";
    
    		return $output;
    	},
    	10,
    	2
    );
  3. Step 3: Load /robots.txt

    Open the address in a browser. WordPress's lines are still there, and yours follow them.

    text
    User-agent: *
    Disallow: /wp-admin/
    Allow: /wp-admin/admin-ajax.php
    
    Sitemap: https://example.com/wp-sitemap.xml
    
    User-agent: *
    Disallow: /example-folder/

The function receives two things: the text WordPress has written so far, and whether the site is open to search engines. Whatever it returns is what a crawler gets, so it must always return the text. Two groups for User-agent: * are not a fault. Google's documentation and the robots.txt standard both say that groups for the same crawler are combined into one.

The file shows on the Plugins screen under "Must-Use" and cannot be switched off there. To undo the change, delete the file. WordPress's documentation notes that must-use plugins get no update notices, so keep a note of why the file exists.

The same add_filter call also works in an ordinary plugin, or in a child theme's functions.php without the opening <?php line and the comment. WordPress loads only the active theme's functions.php, so a rule kept there leaves when the theme is changed.

An SEO plugin's editor

An SEO plugin that offers to edit robots.txt does one of the two things above, and its own documentation says which. Yoast SEO's help page says its file editor creates a robots.txt file and that the file replaces WordPress's default. Rank Math's says its editor changes the answer WordPress writes, and that a real robots.txt file has to be deleted from the root before the editor can be used.

So with any such plugin, look in the site's root first. If a robots.txt file is there, that file is what crawlers get, whatever a settings screen shows.

Which one wins

The file. With a file and a filter both in place, the answer is the file's text alone and the filter's line is absent. The filter only ever changes what WordPress writes, and WordPress is not asked while a file exists.

What robots.txt is for, and what it is not for

robots.txt tells crawlers where they may go. It does not decide what a search engine lists. Google's introduction to robots.txt says the file is mainly for avoiding too many requests to a site, and that "it is not a mechanism for keeping a web page out of Google."

Three consequences, each from Google's documentation:

  • A blocked page can still be listed. If other pages link to an address that robots.txt blocks, Google can list the address without visiting it. The result has no description.
  • A noindex behind a block is never read. For a noindex to work, the page must not be blocked by robots.txt. If it is, the crawler never sees the tag, and the page can still appear in results.
  • It is a request, not a lock. Whether a crawler obeys the file is up to the crawler. Google's advice for anything private is a password.
You wantUse
Crawlers to stay out of part of the siteA Disallow rule in robots.txt
A page to stay out of search resultsA noindex on the page, with the page left open to crawlers
Something to stay privateA password

Two rules that turn up in older advice do nothing for Google: its documentation says a noindex line inside robots.txt is not supported, and neither is crawl-delay.

What not to block on WordPress

  • /wp-admin/admin-ajax.php. WordPress allows it on purpose. If you write your own file and keep Disallow: /wp-admin/, keep the Allow line with it.
  • /wp-includes/ and /wp-content/. On a fresh WordPress the home page loads its stylesheets and scripts from both folders. Google's documentation says not to block resource files a page depends on, because Google does a poor job of analyzing pages it cannot load them for, and that it does not render JavaScript from blocked files.
  • /wp-content/uploads/. That is where your images are. Google's documentation says a robots.txt block on image files keeps them out of Google's search results.
  • A page you have marked noindex. The tag has to be read to work.
  • Everything. Disallow: / under User-agent: * tells every crawler that follows that group to stay out of the whole site. If a site that should be found has that line, see why a WordPress site is not on Google.

How to test a rule

Read the file as it is served. What counts is what a crawler receives, not what is in an editor. Google's documentation suggests opening the address in a private browsing window. From a command line, this prints the status and headers as well as the text:

bash
curl -s -i https://example.com/robots.txt

A working answer starts with a 200 status and a Content-Type of text/plain, followed by your lines. A 301 with a Location header, or a page of HTML, means something other than a robots.txt answered.

Check what Google fetched. Search Console has a robots.txt report among a property's settings. Google's help page says it shows the robots.txt files Google found for the property's hosts, when each was last crawled, whether the fetch succeeded, the file's size and any lines Google could not parse. You can open the last version Google fetched and the versions from the past 30 days, and ask for a recrawl after an important change. The report is offered for Domain properties and for URL-prefix properties with no path. To find out whether one address is blocked, the same help page points to the URL Inspection tool.

Allow for the delay. Google's documentation says it generally caches a robots.txt file for up to 24 hours, so a change is not acted on at once.

AI crawlers and robots.txt

The companies behind AI products publish the names their crawlers answer to in robots.txt. Each line below is what that company's own page says the name is for.

Name in robots.txtWhoseWhat its owner says it is for
GPTBotOpenAICollecting content that may be used to train its models
OAI-SearchBotOpenAIFinding pages to show in ChatGPT's search results
ClaudeBotAnthropicCollecting content that may be used to train its models
Claude-SearchBotAnthropicIndexing content to improve its search results
PerplexityBotPerplexityFinding pages to show and link in Perplexity's results, and not for training models
Google-ExtendedGoogleA name only, with no crawler of its own: whether content Google crawls may be used to train Gemini models and to ground their answers. Google says it has no effect on Google Search.

One rule of the standard matters here. A crawler follows the group that names it, and follows the * group only when no group does. Google's documentation adds that a named group and the * group are not combined. WordPress's default has only a * group, so every crawler that honors robots.txt follows the same three lines. The moment you add a group for one crawler, that crawler stops reading the * group, and the /wp-admin/ lines have to be repeated in its group if you still want them to apply.

A page fetched because a person asked for it is a different case. OpenAI's page says robots.txt rules may not apply to ChatGPT-User, and Perplexity's says Perplexity-User generally ignores them.

Whether to let these crawlers in, and what each choice costs, is a subject of its own. This page covers only where the lines go.

When to get help

  • Ask the host if you cannot find the site's root or cannot upload to it.
  • Hand it over if the address answers with something you cannot trace to a file, a plugin or the filter. A one-time fix from WP Ministry covers one issue on one site and starts with a free diagnosis, which gives you a written cause and a fixed quote.

robots.txt is one of the things WordPress gives a search engine with no plugin installed. WordPress SEO without a plugin lists the rest.

Common questions

Where is the robots.txt file in WordPress?

On a site where nobody has added one, there is no file. WordPress writes the answer each time /robots.txt is requested. If a file exists, it is in the site's root, the folder that holds wp-admin, wp-content and wp-includes, and it is served in place of WordPress's answer.

Does a WordPress site need a robots.txt file?

No. WordPress answers the address without one. Google's help for the robots.txt report also says that having no robots.txt at all is fine, and means Google can crawl every address on the site.

Does Disallow: /wp-admin/ hide my login page?

No, for two reasons. The login page is wp-login.php in the site's root, not inside /wp-admin/, so the rule does not match it. And a robots.txt rule is a request to crawlers, not a barrier: anyone can still open the address.

I changed robots.txt. Why does Google still show the old one?

Google's documentation says it generally keeps a copy of the file for up to 24 hours. The robots.txt report in Search Console shows the version Google last fetched and lets you ask for a recrawl.

Can I put a robots.txt in a subfolder for a WordPress installed there?

Not one that crawlers will read. Google's documentation says the file has to be at the root of the host and that crawlers do not check subdirectories. For WordPress at example.com/blog/, the file that applies is example.com/robots.txt, and you write it by hand.

More on this subject

Quick Fix, done for you

Quick Fix is $49. One issue, one site, up to about an hour. No fix, no fee. 30-day warranty. It starts with a free diagnosis.