Skip to content

Staging site indexed by Google: how to remove it and keep the next copy out

Put a password on the staging site, or a noindex that Google is allowed to read, then ask for removal in Search Console's Removals tool, which lasts about six months. Do not rely on Disallow in robots.txt. Google says not to, and it stops a noindex from being read.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • A staging copy is a whole site with indexing switched on. A copy of the live database brings the live site's search engine setting with it.
  • A password is the fix that lasts. A request without it is answered 401, and Google removes addresses that answer that way.
  • Where a password is not possible, tick "Discourage search engines from indexing this site" and leave robots.txt alone, so the noindex can be read.
  • The Removals tool in Search Console hides the addresses for about six months. It is the quick step, not the lasting one.
  • A Disallow rule in robots.txt is the one method Google's removal instructions say not to use.
  • Tick the box again after every import of the live database, and keep the box and the password lines off the live site.

A staging site turns up in Google because it is a complete copy of the live site, open to anyone, with nothing on it that tells a search engine to stay away. To get it out, make the copy answer with a password, then ask Google to remove its addresses with the Removals tool in Search Console. Where a password is not possible, use a noindex that Google is allowed to read.

Do not rely on Disallow: / in robots.txt. Google's removal instructions say not to use robots.txt for this, and a robots.txt block stops Google from reading a noindex at all. The order matters too: Google's removal lasts about six months, so whatever keeps the copy out for good goes in first.

How a staging copy gets into Google

Google's documentation says how it learns that an address exists: it follows a link from a page it already knows, or it reads a sitemap. Nothing in that process asks whether a site is meant to be public. One link to staging.example.com from a page Google has visited is enough.

A copy also carries the live site's answer to search engines. The box "Discourage search engines from indexing this site" is stored in the database as the option blog_public, and a copy of the live database is a copy of that row. On a live site that is meant to be found the box is not ticked, so on the copy it is not ticked either. With WP-CLI, in the copy's folder:

bash
wp option get blog_public

It prints 1 on a site that is open to search engines, a fresh WordPress included, and 0 when the box is ticked. A site at 1 prints this robots tag on every page, which blocks nothing:

text
<meta name='robots' content='max-image-preview:large' />

It also answers at /wp-sitemap.xml and names that sitemap in its robots.txt, as the live site does.

Google then has the same pages at two addresses. Its page on canonicalization says it groups pages with the same content and picks one to show, and that it "may choose a different page as canonical than you do". That is how a staging address can appear where the real one should.

Step 1: put a password on the copy

A password is the fix that lasts. Google's documentation says that password-protecting content prevents it from appearing in Google Search, and that if it already appears, Google "will eventually remove that content from our search results". A request without the password is answered with status 401, and Google's page on status codes says that addresses in its index that answer with a 4xx status are removed from it.

Look first in your hosting control panel for a way to password-protect a folder or a subdomain. If it has one, use it and go to the two checks at the end of this section.

On an Apache server you can do it by hand with HTTP Basic authentication. This needs SSH access, and a server that accepts authentication settings in .htaccess (Apache's documentation names the setting that allows it, AllowOverride AuthConfig). If the server refuses the lines, take them out and ask the host.

  1. Step 1: Create the password file

    preview is the user name you will type later. The path is where the file is kept. Apache's documentation says to put it somewhere not accessible from the web, so use a folder outside the one the site is served from. The command asks for a password twice and answers Adding password for user preview.

    -c creates the file, so leave it off when adding a second user to the same file. -B stores the password as a bcrypt hash and not as text.

    bash
    htpasswd -c -B /home/example/.htpasswd preview
  2. Step 2: Add four lines to .htaccess

    Open the .htaccess file in the staging site's folder and put these at the top, above the line # BEGIN WordPress. The path is the file you just made.

    .htaccess
    AuthType Basic
    AuthName "Staging"
    AuthUserFile "/home/example/.htpasswd"
    Require valid-user
  3. Step 3: Ask for a page without the password

    It prints 401. That is what a crawler now receives for every address on the copy: pages, the login screen, images and robots.txt alike.

    bash
    curl -s -o /dev/null -w "%{http_code}\n" https://staging.example.com/
  4. Step 4: Ask again with it

    Put your own password in place of your-password.

    It prints 200. In a browser, the address now asks for the name and password before it shows anything.

    bash
    curl -s -o /dev/null -w "%{http_code}\n" -u "preview:your-password" https://staging.example.com/

Apache's documentation warns that Basic authentication sends the password unencrypted, so use the copy over HTTPS. And be sure the .htaccess you edit is the copy's. The same four lines on the live site answer 401 to every visitor and every crawler.

Where a password is not possible: a noindex Google can read

Some copies cannot sit behind a password, because an outside service has to reach them. Then tick "Discourage search engines from indexing this site" on the copy, under Settings, then Reading. With WP-CLI:

bash
wp option update blog_public 0

Then read what a page serves, with the copy's address in place of https://staging.example.com/:

bash
curl -sL https://staging.example.com/ | grep -i -o "<meta[^>]*robots[^>]*>"
text
<meta name='robots' content='noindex, nofollow' />

Google's documentation says that when its crawler fetches a page and reads noindex, it drops that page from its results, whether or not other sites link to it. The same page gives the condition: the page "must not be blocked by a robots.txt file". WordPress keeps to it. With the box ticked, this is everything its robots.txt says:

text
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Crawlers are kept out of the dashboard and nothing else, so every page can be fetched and its tag read. Leave it that way.

This is the weaker measure. The copy still opens for anyone who has the address, and Google's Removals help calls a noindex "less secure than the other methods". What the box does has the rest, including the files the tag does not cover.

Step 2: ask Google to remove the addresses

The password or the tag takes effect as Google fetches each address again, and Google gives no date for that. The Removals tool is the quick part. Its help page says a request "usually takes up to a day to process, and is not guaranteed to be accepted".

The tool is in Search Console, and the address has to be in a property you own. Google's help says a Domain property for example.com includes all its subdomains and is verified with a DNS record. A URL-prefix property covers only addresses that begin with its prefix, so one made for the live site does not include staging.example.com.

  1. Step 1: Open the tool and start a request

    In Search Console, open Removals. On the "Temporary Removals" tab, select "New Request".

  2. Step 2: Name the whole copy

    Choose "Temporarily remove URL", then "Remove all URLs with this prefix", and enter the copy's address with its closing slash, such as https://staging.example.com/.

  3. Step 3: Send it and look again later

    Select "Next". The request is listed on the same tab with its status.

Three things on that help page matter for a staging copy.

  • The live site is not touched. In Google's words, "Blocking https://staging.domain will not block www.domain or https://domain.com." The reverse is not true: a request for the bare domain covers its www, non-www, http and https forms, which is the live site. If the copy lives in a folder, such as example.com/staging/, enter that folder with its closing slash and read it twice before sending.
  • It runs out. "A successful request lasts only about six months." The page adds that a block does not stop crawling, and that a page crawled before it is removed or password-protected "can appear in search results after your temporary blackout expires". That is why step 1 comes first.
  • It is Google only. The tool "removes content only from Google Search". The password is what covers every other search engine.

What not to do

A Disallow rule in robots.txt and nothing else

This is the usual advice, and it is the one method Google's instructions rule out. The Removals help lists the ways to make a removal permanent and ends the list with "Do not use robots.txt as a blocking mechanism." Google's introduction to robots.txt gives the reason: the file "is not a mechanism for keeping a web page out of Google", and a disallowed page "can still be indexed if linked to from other sites".

A block also undoes a noindex, because the crawler never fetches the page that carries it. WordPress's developer note for version 5.3 gives the same reason for no longer writing Disallow: / when the box is ticked. WordPress robots.txt covers the file itself.

Deleting the copy without checking what answers

Removing the copy is a sound choice. Google's page on removals calls removing content the most secure way. But the Removals help asks for more than deleted files: the server has to answer 404 (Not Found) or 410 (Gone).

What a deleted copy answers depends on the host. An emptied folder may show a placeholder page, and a staging name that now leads to the live site's folder serves the live site a second time. Run the first status check from step 1 against the old address. A 200 means the address still serves something.

If the staging name stays on the server, these two lines as the whole of its .htaccess answer 410 for every address:

.htaccess
RewriteEngine On
RewriteRule ^ - [G]

If the name is taken out of DNS instead, nothing answers. Google's documentation says it treats DNS errors much as it treats server errors, and that "already indexed URLs that are unreachable will be removed from Google's index within days".

A canonical tag as the only measure

A plugin can make each page of the copy name the live page as its canonical address. Google's documentation says that indicating a canonical preference "is a hint, not a rule". It also hides nothing: the copy still opens for anyone. Without such a plugin, WordPress's own canonical tag on a copy names the copy's address.

Keeping the next copy out

  • Password first. Put the password on the staging address before the copy is made, or straight after. If the live site's files are copied into the folder, its .htaccess replaces yours. Add the four lines again and run both status checks.
  • Tick the box after every import. Each import of the live database sets blog_public back to the live site's 1. Run wp option update blog_public 0 on the copy after each one.
  • Check what a host's staging tool does. It may set a password, tick the box, or do neither. Do not assume. Run the checks below on the copy's address. How to set up a WordPress staging site covers what else a copy needs, such as stopping its email.
  • Mind the way back. Two things must not travel from the copy to the live site: the ticked box, which arrives with a pushed database, and the password lines, which arrive with a pushed .htaccess. Either one takes the live site out of search results as Google crawls it. Turning the box off and the WordPress launch checklist cover the first.

Checking that it worked

  • The status. The first status check prints 401 for a copy behind a password, and 404 or 410 for one that is gone.
  • The tag. Where there is no password, the tag check prints noindex.
  • From a browser. The page indexing check reads the same things from an address you give it: the status it answers with, any rule in robots.txt, and a noindex on the page or in its headers.
  • In Google. Search for site:staging.example.com. Results mean the copy is still listed. No results is a good sign and not proof: Google's page on the operator says the list "is not always exhaustive" and that the operator "doesn't necessarily return all the URLs that are indexed". For one address, the URL Inspection tool in Search Console gives Google's own answer.

When to get help

  • Ask the host if the server is not Apache, or if it will not accept the lines in .htaccess.
  • Hand it over if a launch or a rebuild is close. SEO for a new website from WP Ministry is a one-time setup for a new or rebuilt WordPress site, and making sure the staging copy stays out is part of it.

Common questions

Will removing the staging site from Google affect my live site?

Not if the copy has its own subdomain. The password is on the staging address only, and Google's help says a removal request for a staging subdomain does not block the www or bare domain. Take care when the copy is a folder of the live site: enter that folder, never the domain alone.

How long does it take for the staging pages to go?

Google's help says a removal request usually takes up to a day to process and is not guaranteed to be accepted. That hides the addresses for about six months. The lasting removal happens as Google fetches each address again and meets the password or the noindex, and Google gives no date for it.

Is ticking "Discourage search engines" enough for a staging site?

It is a request to search engines and it keeps nobody out. Google drops pages that carry the tag once it has crawled them, but the copy still opens for anyone with the address. Use it together with a password, or instead of one only when a password is not possible.

Should I add Disallow to robots.txt as well, to be safe?

No. A crawler that obeys the block never fetches the pages, so it never sees the noindex on them, and Google can still list a blocked address that other sites link to.

Can I just delete the staging site?

Yes. Google's page on removals calls removing the content the most secure way. Afterwards, ask for one of its old addresses and read the status. It should be 404 or 410, or no answer at all. If a page still comes back, the name is serving something.

More on this subject

SEO setup for a new site, done for you

SEO setup for a new site is $299. Indexing, sitemap, Search Console and Bing Webmaster Tools set up and checked on a new or rebuilt site. It starts with a free diagnosis.