Skip to content

WordPress search results pages (/?s=) in Google and Search Console: what they are and what to do

An address with ?s= or /search/ in it is your site's own search page, asked for with someone else's words. WordPress puts noindex on every search page, so Google reads it and leaves it out. Check that your site prints the tag. Nothing was added to the site.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • An address with ?s= or /search/ in it is WordPress's search results page. Anyone can make one by typing an address, and nothing is saved on the site.
  • WordPress prints a robots tag of noindex, follow on every search results page, with results or without, and has since version 5.7.
  • A search that finds nothing still answers 200. The searched words come back in the title and the heading as text, never as markup or a link.
  • Search addresses are not in wp-sitemap.xml, and WordPress's robots.txt says nothing about them.
  • In Search Console they belong under "URL marked ‘noindex’". For a page you do not want indexed, Google's help calls that the outcome to be glad of.
  • Spam words on an address that is not a search, or a page of links you never wrote, is a different problem. That one is a break-in.

An address on your site with ?s= or /search/ in it is WordPress's own search results page. Nobody put a page on your site. Someone asked your site to search for their words, and WordPress answered with a page that repeats those words in its title, as it does for any search. A link to that address from somewhere else is how a search engine came to know of it.

WordPress already asks search engines to leave these pages out. Every search results page carries a robots tag that says noindex, follow, whether the search found anything or not. Confirm that your site prints the tag. If it does, a list of these addresses in Search Console is Google recording that it read the tag, and there is nothing to clean up.

What these addresses are

WordPress answers a search at two forms of address.

FormExampleWhen it answers
The query formhttps://example.com/?s=oak+tableAlways
The path formhttps://example.com/search/oak+table/When permalinks are set to anything but Plain

The search box on a site sends visitors to the first form, and neither form needs the box. Anyone can type an address with any words after ?s=, and WordPress runs the search. Nothing is stored. The page exists only as the answer to that request.

That is what a spammer uses. The words in the address come back in the page's title and heading, on your domain, and a link from another site is enough for a crawler to ask for it. The change that added noindex to search results in WordPress 5.7 names the problem in its own description: "reflected web spam attacks".

What WordPress answers

These are the answers of a new WordPress with the Twenty Twenty-Five theme, no plugin, Post name permalinks and one post, "Hello world!".

Address requestedStatusRobots tagTitle begins
/?s=hello, which finds the post200noindex, follow, max-image-preview:largeSearch Results for “hello”
/?s=buy+cheap+example+pills, which finds nothing200The sameSearch Results for “buy cheap example pills”
/search/hello/200The sameSearch Results for “hello”
/search/buy+cheap+example+pills/200The sameSearch Results for “buy cheap example pills”
  • The tag is on all four. Google's documentation gives noindex this meaning: "Do not show this page, media, or resource in search results." follow is not a rule on Google's list. Following links is what it does unless a page says nofollow.
  • A search that finds nothing answers 200. The page says "Sorry, but nothing was found." WordPress's source lists a search among the requests it does not turn into a 404 when no post matches.
  • The words come back as text. They are in the title, in the heading ("Search results for:") and in the search box. A search for <script>alert(1)</script> came back as &lt;script&gt;alert(1)&lt;/script&gt;, which a browser shows and does not run, and a search holding an HTML link came back the same way, as characters and not a link.
  • None of them is in the sitemap. WordPress's developer note says its sitemap covers public post types and taxonomies, author archives and the home page. No file under /wp-sitemap.xml on the test site held a search address after the searches above had been made.
  • robots.txt says nothing about them. WordPress's answer at /robots.txt keeps crawlers out of /wp-admin/ and nothing else. The section on robots.txt below says why that is right.

One thing is not covered by the tag. Each search page links, in its <head>, to an RSS feed of the same search, at /search/hello/feed/rss2/. The feed answered 200 with the searched words in its title. A feed has no place for a robots tag, and WordPress sent no X-Robots-Tag header with it.

Check what your own site prints

This asks your site for a search and prints the robots tag and the title of the page that comes back. Put your own address in place of https://example.com.

bash
curl -s "https://example.com/?s=test" | grep -Eio "<title>[^<]*</title>|<meta name=.robots.[^>]*>"
text
<meta name='robots' content='noindex, follow, max-image-preview:large' />
<title>Search Results for &#8220;test&#8221; &#8211; Example</title>

The word to look for is noindex, in the first line. Without a terminal, search for anything on your site, view the page's source and look for the same tag.

If the site's permalinks are not Plain, run it once more with /search/test/ in place of /?s=test. If the line has no noindex, or there is no robots line at all, go to "What can take the tag away" below.

Why Google lists them in Search Console

Google's help for the Page indexing report says Google "can find a page in many different ways, including someone linking to your page from another site." That is how it finds a search address you never linked to. It asks for the page, reads the tag and files the address among the pages that are not indexed, with a reason.

For a page that carries the tag, the reason in the help page's list is "URL marked ‘noindex’". Its description reads: "When Google tried to index the page it encountered a 'noindex' directive and therefore did not index it. If you do not want this page indexed, congratulations!" The same page says addresses that are not indexed "can be fine", and gives a noindex tag as one of the right reasons.

So hundreds of search addresses under that reason are the tag doing its work.

If a search address shows in Google's results themselves, Google's documentation gives three explanations:

  • Google has not crawled it since the tag was there. It has to crawl a page to see the tag, and "it may take months for Googlebot to revisit a page". A site on a WordPress older than 5.7 printed no such tag.
  • robots.txt blocks the address, so the tag is never read. The next section covers this.
  • The tag is missing. Run the check above.

For one result that has to go sooner, the URL Inspection tool can ask for a recrawl, and the Removals tool hides an address, or every address that starts with a prefix, for about six months. Google's help says the tool hides a result and does not stop crawling. The tag is what lasts.

Why a Disallow in robots.txt is not the fix

Adding Disallow: /?s= to robots.txt looks like the obvious fix. It stops the crawling and gives up the tag. Google's page on noindex says that for the rule to work the page "must not be blocked by a robots.txt file", because then "the crawler will never see the noindex rule". Its introduction to robots.txt adds that a blocked address "can still appear in search results" when other pages link to it. Links from other pages are how these addresses arrive.

Leave crawling open, as WordPress doesBlock search addresses in robots.txt
A crawler that obeys robots.txtAsks for the page and reads noindexDoes not ask for the page
What Google can do with the addressLeaves it out of resultsCan still list the bare address, with no description
Requests to your serverOne for each crawl of each addressNone from crawlers that obey the file
Reason in Search Console"URL marked ‘noindex’""URL blocked by robots.txt", or "Indexed, though blocked by robots.txt"

There is one case for blocking anyway: a flood large enough that the crawling itself strains the server. Google's page on URL structure says to consider robots.txt for "dynamic URLs, such as URLs that generate search results" when crawling of them is the problem. That is a choice about load, with the cost in the table. These are the lines:

text
User-agent: *
Disallow: /?s=
Disallow: /search/

A rule matches every address that starts with it, so the first line misses an address where s= is not the first parameter. Where robots.txt is in WordPress and how to change it has the two ways to add lines, and the robots.txt tester shows whether a rule matches an address.

Make a search that finds nothing answer 404

This step is optional. The tag already covers every search page. What it adds is a status code that says what the page says in words, and it reaches the feed of such a search, which the tag does not.

Google's page on status codes says that when a page answers 200 and what it holds is "an empty page or an error message", Search Console shows a soft 404 error. It also says Google does not index an address that answers with a 4xx code, and ignores what such a page holds. A search for words that are in none of your posts finds nothing, so a 404 for an empty search covers those addresses without relying on a tag.

  1. Step 1: Add the file

    Look in wp-content. If there is no folder named mu-plugins, create one. WordPress loads every PHP file in that folder on each request, with nothing to activate. Save this in it.

    wp-content/mu-plugins/empty-search-404.php
    <?php
    /**
     * Plugin Name: 404 for a search that finds nothing
     * Description: A search results page with no results answers 404 in place of 200.
     */
    
    add_action(
    	'wp',
    	function () {
    		global $wp_query;
    
    		if ( is_search() && ! is_admin() && 0 === $wp_query->post_count ) {
    			status_header( 404 );
    			nocache_headers();
    		}
    	}
    );
  2. Step 2: Ask for a search that finds nothing

    This prints the status code alone. It printed 200 before the file was added and 404 after.

    bash
    curl -s -o /dev/null -w "%{http_code}\n" "https://example.com/?s=zx81-nothing-matches-this"

What it does to visitors: nothing they can see. On the test site the page was the same "nothing was found" page, with the same title, heading and noindex tag. A search that found the post still answered 200, in both forms, and so did a search in the dashboard. The feed of an empty search answered 404 as well. Anything on the site that logs 404s will now log empty searches. To undo it, delete the file.

When it is something worse

A search address is harmless because nothing is on the site. Google's spam policies define hacked content as "any content placed on a site without permission, due to vulnerabilities in a site's security", and a search address places nothing. The same policies have no entry for a site's own search being used this way.

These signs point to a break-in instead:

  • The spam is at an address that is not a search. No ?s= and no /search/: a page, a post or a folder you never made.
  • The page holds content. A search page for spam words says nothing was found. A page of products, foreign text or links to other sites was written by someone.
  • You are sent somewhere else when you open the result from Google.
  • Search Console's Security issues report has an entry. Google's help says the report shows its findings when it determines that a site was hacked.

For results in Japanese, or for pills, loans or gambling, at addresses of that kind, see spam pages in Google and how to get them out. The hacked-site check asks for one page twice, once plainly and once as a visitor arriving from a search result, and reports signs such as a redirect to another site or links hidden in the page. How to check whether a WordPress site has been hacked covers the checks from inside.

What can take the tag away

  • A WordPress older than 5.7. The function that adds the tag was introduced in that version.
  • Code on the wp_robots filter. WordPress builds the tag through one filter, and a theme, a plugin or a snippet can take noindex out of it. With WordPress's own function unhooked on the test site, the check above printed max-image-preview:large and no noindex.
  • A search that does not go through WordPress. A plugin or a theme that serves results at an address of its own, or builds pages for searched terms, is outside this tag. Read what those addresses print with the same check.

An SEO plugin can print a robots tag of its own, so go by what the check prints. Yoast SEO's documentation says the plugin puts noindex on search pages automatically. It also describes settings named "Internal site search cleanup", which can filter searches over a length or with common spam patterns, redirect the path form to the query form, remove search feeds, and add a robots.txt rule against crawling search addresses. Of the last, its own page says blocking is not its general advice and is for search pages that are "under attack" or crawled to excess.

What WordPress prints with no plugin is in WordPress SEO without a plugin. Two neighbors of this problem have their own pages: attachment pages in Google, and the box that discourages search engines, which changes this tag to noindex, nofollow.

When to get help

Get help if the check prints no noindex and you cannot find what removes it, or if Search Console lists many kinds of address you do not recognize and you want to know which matter. A WordPress SEO audit from WP Ministry is a one-time job: we look at what a search engine meets when it comes to your site, write down what is in its way, fix what can be fixed on the WordPress side, and test each fix again. If the signs point to a break-in, start with the pages on hacked sites above.

Common questions

Has my site been hacked if Google lists spam search addresses?

Not from that alone. A search address is made by whoever types it, and nothing is saved on your site. Open one: if it is your theme's search page saying nothing was found, with the spam words only in the title and heading, it is this problem. Spam at addresses that are not searches is the other one.

Should I add Disallow: /?s= to robots.txt?

Not to get the addresses out of Google. A blocked page is never fetched, so its noindex is never read, and Google's documentation says a blocked address can still be listed. Blocking is for a site where the crawling itself is a load problem, and it trades the tag away.

How long until they leave the Search Console report?

Google gives no fixed time. An address under "URL marked ‘noindex’" is already out of the results, and the entry is a record of that. Google's documentation says it may take months for a page to be crawled again, and nobody outside Google sets that date.

More on this subject

SEO audit, done for you

SEO audit is $249. A written audit, each finding with its source, and up to 2 hours of fixes done and tested again. It starts with a free diagnosis.