Skip to content

WooCommerce filter URLs indexed in Google: ?orderby, ?filter_ and add-to-cart addresses, and what to do

A WooCommerce store answers at extra addresses for one list of products, such as ?orderby=, ?filter_, ?min_price=, /page/2/ and ?add-to-cart=. Each returns the page with no canonical link and no noindex. A small store can leave them. To stop the crawling, add robots.txt rules for the parameters.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, WooCommerce 11.2.0, PHP 8.3.35

In short

  • A sorted or filtered shop address answers 200 with the plain page's robots tag and no canonical link. WordPress prints a canonical link on single pages only, and WooCommerce adds none to a list.
  • WooCommerce already asks crawlers to stay away from add-to-cart addresses, in the robots.txt that WordPress answers with. A real robots.txt file drops those lines.
  • The cart, checkout and my-account pages carry noindex on purpose. Never block them in robots.txt.
  • For filter addresses nobody needs in search, Google's first answer is a robots.txt rule. It calls a canonical link and rel="nofollow" less effective in the long term.
  • Most of the statuses these addresses sit under in Search Console are not errors and need nothing done.
  • Do not put noindex on an address that robots.txt blocks, and do not redirect filter addresses to the plain one.

The addresses with question marks in your page indexing report are your shop and category pages, sorted or filtered. WooCommerce makes them from parameters: orderby for the sort order, names that begin filter_ and query_type_ for attribute filters, min_price and max_price, rating_filter, and add-to-cart. Each one answers with status 200 and the same products in another order, or fewer of them. On a store with no SEO plugin, none of them carries a canonical link or a noindex.

For most stores the report is describing something that works as intended, and nothing needs fixing. When the addresses run to thousands, or crawlers spend the server's time on them, the remedy Google names first is a robots.txt rule that keeps crawlers out. This page lists the addresses, shows what each one tells a search engine, and gives the rules.

The addresses one list of products answers at

These were read from a test store with twenty products, one category named Mugs and one attribute named Color, on WordPress's default theme. The plain addresses are /shop/ and /product-category/mugs/.

ParameterExampleWhat it doesWhat makes a link to it
orderby/shop/?orderby=priceSorts the list. The values on offer are popularity, rating, date, price and price-desc.The sort drop-down, which is a form and not a link
filter_<attribute> with query_type_<attribute>/shop/?filter_color=red&query_type_color=orKeeps the products with that value. query_type_ says whether several values mean any of them or all of them.An attribute filter
min_price, max_price/shop/?min_price=5&max_price=10Keeps the products in a price rangeA price filter, which is a form or a slider
rating_filter/shop/?rating_filter=5Keeps the products with that average ratingA rating filter
filter_stock_status/shop/?filter_stock_status=instockKeeps the products with that stock statusThe Product Filters block
categories, tags, brands/shop/?categories=mugsKeeps the products in a category, a tag or a brandThe Product Filters block
A page number in the path/shop/page/2/Shows the next page of the listThe page links under the list
add-to-cart/shop/?add-to-cart=10Puts the product with ID 10 in a cart, then shows the pageThe add to cart control

Every one of them works on a category page too: /product-category/mugs/?orderby=price is the category, sorted. A product's own page has a further family, the attribute_ parameters that choose a variation, and WooCommerce variations and search covers those.

Almost none. On the default theme, the only links on the shop page that carry a parameter or a page number are the page links. The sort control is a drop-down inside a form. The add to cart control is a button with no address in it. There is no filter on the page until someone adds one.

That matters because of how Google finds addresses. Its documentation on pagination says it generally crawls the addresses found in the href of <a> elements, and that its crawlers do not click buttons.

Filters come in two kinds, and only the older kind prints links:

  • The classic widgets. "Filter Products by Attribute" prints each value as a link, and marks each link rel="nofollow". "Filter Products by Rating" prints each rating as a link with no rel at all. "Filter Products by Price" is a form. "Active Product Filters" prints a link for removing each filter, marked rel="nofollow".
  • The Product Filters block. It prints checkboxes and buttons, and changes the address with a script when one is pressed. Its choices are not links in the page.

The add to cart control becomes a link in two cases: a theme or a shortcode that uses WooCommerce's classic product list, and a store where "Enable AJAX add to cart buttons on archives" is unticked under WooCommerce, then Settings, then Products. In both, the link is ?add-to-cart= and the product's ID, marked rel="nofollow".

One parameter leads to more

WordPress's page links keep whatever follows the question mark. On /shop/?orderby=price-desc, the link to the second page is /shop/page/2/?orderby=price-desc, and it carries no rel="nofollow".

That holds for a parameter WooCommerce has never heard of. /shop/?anything=1 answers 200 with the full list, and links to /shop/page/2/?anything=1. So a crawler that meets one such address, on your site or in a link from someone else's, is handed more of them. Google's documentation on link attributes makes the same point from the other side: a page behind a nofollow link may still be found through other means, such as a link from another site.

What each address tells a search engine

This is what WooCommerce and WordPress print by themselves, with no SEO plugin, as read from the test store.

AddressStatusCanonical linkRobots tag
/shop/200Nonemax-image-preview:large
/shop/?orderby=price200Nonemax-image-preview:large
/shop/?filter_color=red200Nonemax-image-preview:large
/shop/?min_price=5&max_price=10200Nonemax-image-preview:large
/shop/?rating_filter=4, which matches no product200, with "No results found"Nonemax-image-preview:large
/shop/page/2/200Nonemax-image-preview:large
/shop/page/3/, past the last page404Nonemax-image-preview:large
/product-category/mugs/?orderby=price200Nonemax-image-preview:large
/shop/?add-to-cart=10200, and a cart is madeNonemax-image-preview:large
/product/mug-1/?add-to-cart=10200, and a cart is madeThe product's plain addressmax-image-preview:large
/?s=mug&post_type=product, a product search200Nonenoindex, follow, max-image-preview:large
/cart/200Itselfmax-image-preview:large, noindex, follow
/checkout/ with an empty cart302, to /cart/
/checkout/ with a product in the cart200Itselfmax-image-preview:large, noindex, follow
/my-account/200Itselfmax-image-preview:large, noindex, follow

Four things to take from it:

  • The robots tag blocks nothing. max-image-preview:large is on every WordPress page and allows large image previews. The word that keeps a page out of search is noindex, and no sorted or filtered address has it. WordPress itself puts it on search results, which is why a product search does.
  • A list has no canonical link. WordPress prints one for a single post, page or product: its rel_canonical() function stops at once for anything else. A shop page and a category page are lists, so neither gets one, with a parameter or without, and WooCommerce adds none. A product does, which is why an add-to-cart address on a product page points back at the plain product.
  • A filter that matches nothing is still a page. It answers 200 and says "No results found". A page number past the end answers 404.
  • An add-to-cart address is not a view of a page. Each request makes a cart: WooCommerce sets its cart cookies and saves a session in the database. It is the one address here that changes something on the server every time a crawler asks for it.

A theme, an SEO plugin or a filter plugin changes this table, so read your own store's. This prints the canonical link and the robots tag of one address, as a crawler receives them:

bash
curl -s "https://example.com/shop/?orderby=price" | grep -i -o -E '<(link|meta)[^>]*(canonical|robots)[^>]*>'

On the test store it prints one line, the robots tag, and no canonical link:

text
<meta name='robots' content='max-image-preview:large' />

An SEO plugin can fill the canonical column. Yoast SEO's published specification, for one, says a canonical link belongs on list pages as well, with sorting and filtering queries taken off it and the page number kept. On a store like that, a sorted address points at the plain one.

The cart, checkout and account pages

WooCommerce puts noindex on the cart, checkout and my-account pages on purpose: its function wc_page_no_robots adds the word through WordPress's own robots tag, and the table shows it on all three. When Search Console lists them under a noindex status, the store is working as designed and there is nothing to fix.

What Google says about filter addresses

Google documents this under the name faceted navigation. Its account of the problem is that filters built on URL parameters can make a near-endless number of addresses. A crawler cannot tell whether one is useful without fetching it, so it fetches a very large number of them before it works out that they are not. That leaves less time for the new addresses that are useful.

It then sorts sites into two. A site that does not need its filter addresses in search should prevent them being crawled. A site that does need them should follow a short list of practices, and accept the cost in server resources.

To prevent crawling, it gives these ways, in this order:

Google's wayWhat Google says of itOn a WooCommerce store
robots.txt rules that disallow the parametersThere is often no good reason to let filtered lists be crawled. Allow the items' own pages and one listing page with no filter.WooCommerce writes rules for add-to-cart and for nothing else. The rest are yours to add.
URL fragments for filters, the part of an address after #Google Search generally does not use fragments in crawling or indexing, so filters kept there have no effect on crawling.WooCommerce's filters use parameters. No setting moves them to fragments.
A canonical link to the unfiltered addressMay, over time, decrease how much the other versions are crawled. Generally less effective in the long term than the two above.No list page has a canonical link unless a plugin adds one.
rel="nofollow" on links to filtered pagesMay help, but every link to an address has to carry it. Also generally less effective in the long term.Only some links carry it: the attribute widget's and the add to cart link do, and the rating widget's and the page links do not.

For a site that wants its filter addresses in search, Google asks that a filter combination with no results answer 404, and a page number that does not exist too. WooCommerce does the second and not the first, as the table above shows.

What you can do

Leave it alone when the numbers are small

Google's guide to crawl budget says who it is written for: sites with more than a million pages that change weekly, sites with more than ten thousand that change daily, and sites where a large share of all addresses sit under "Discovered - currently not indexed". It calls those figures rough estimates. It also says that if your pages are crawled the day they are published, you do not need the guide.

Its documentation on canonical addresses says a site will likely do fine without stating a preference, because Google then chooses the version to show. Its help for the Page indexing report says not to expect every address on a site to be indexed, only the canonical ones.

So a store with a few hundred products, whose report lists a few hundred sorted and filtered addresses under the statuses in the table further down, is in the state Google describes as normal. Act when one of these is true: a large share of the store's addresses are "Discovered - currently not indexed", new products wait a long time to be crawled, or the server's logs show crawlers spending their visits on filter addresses.

See what your robots.txt already says

bash
curl -s https://example.com/robots.txt

On the test store, with nothing added, the answer is this:

text
User-agent: *
Disallow: /wp-content/uploads/wc-logs/
Disallow: /wp-content/uploads/woocommerce_transient_files/
Disallow: /wp-content/uploads/woocommerce_uploads/
Disallow: /*?add-to-cart=
Disallow: /*?*add-to-cart=
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml

The five lines after User-agent: * are WooCommerce's. It adds them to the answer WordPress writes when the site has no robots.txt file, and two of them ask crawlers to leave every add-to-cart address alone. Nothing in this answer mentions orderby or a filter.

If your store's answer has no add-to-cart lines, a real robots.txt file or a plugin's editor has replaced WordPress's answer, and WooCommerce's lines went with it. Where robots.txt is in WordPress explains which of the two you have.

Add rules for the sort and filter parameters

A filter on WordPress's robots.txt answer keeps WordPress's lines and WooCommerce's, and adds yours. The smallest place for it is a must-use plugin: one PHP file in wp-content/mu-plugins/, which WordPress loads on every request with nothing to activate.

  1. Step 1: Create the mu-plugins folder if it is not there

    Look in wp-content. If there is no folder named mu-plugins, create one.

  2. Step 2: Add the file

    Save this as store-robots-rules.php in that folder. It puts six lines into the group that WordPress and WooCommerce already wrote.

    wp-content/mu-plugins/store-robots-rules.php
    <?php
    /**
     * Plugin Name: Store robots.txt rules
     * Description: Asks crawlers to stay out of the sorted and filtered copies of the shop's pages.
     */
    
    add_filter(
    	'robots_txt',
    	function ( $output ) {
    		$rules = array(
    			'Disallow: /*?*orderby=',
    			'Disallow: /*?*filter_',
    			'Disallow: /*?*query_type_',
    			'Disallow: /*?*min_price=',
    			'Disallow: /*?*max_price=',
    			'Disallow: /*?*rating_filter=',
    		);
    
    		return preg_replace( '/^User-agent: \*\s*$/m', "User-agent: *\n" . implode( "\n", $rules ), $output, 1 );
    	},
    	20
    );
  3. Step 3: Load /robots.txt

    Run the command from the last section again, or open the address in a browser. The six lines sit directly under User-agent: *, above WooCommerce's.

    text
    User-agent: *
    Disallow: /*?*orderby=
    Disallow: /*?*filter_
    Disallow: /*?*query_type_
    Disallow: /*?*min_price=
    Disallow: /*?*max_price=
    Disallow: /*?*rating_filter=
    Disallow: /wp-content/uploads/wc-logs/
    Disallow: /wp-content/uploads/woocommerce_transient_files/
    Disallow: /wp-content/uploads/woocommerce_uploads/
    Disallow: /*?add-to-cart=
    Disallow: /*?*add-to-cart=
    Disallow: /wp-admin/
    Allow: /wp-admin/admin-ajax.php
    
    Sitemap: https://example.com/wp-sitemap.xml
  4. Step 4: Test the rules against real addresses

    Copy the whole answer into the robots.txt tester, which works in your browser. Give it /shop/, a product's address and /shop/page/2/: each should come back allowed. Then give it a sorted address and a filtered one from your report: each should come back blocked, with the line that decides.

In Google's reading of robots.txt, * stands for any run of characters. /*?*orderby= therefore matches any address that has orderby= somewhere after its question mark, whatever comes before it. Google's own example for faceted navigation is written the same way. filter_ and query_type_ have no = because the attribute's name follows them, and filter_ also covers filter_stock_status.

The rules leave three of the block's parameters out: categories, tags and brands. Those words are common enough in other plugins' addresses that a rule for them should be a decision. Add a line in the same form if your report shows them.

The file is listed on the Plugins screen under "Must-Use" and cannot be switched off there. To undo the change, delete the file.

If the site has a real robots.txt file

The filter does nothing then, because WordPress is never asked. Put the rules in the file itself, with the lines WordPress and WooCommerce would have written, since the file is the whole answer.

robots.txt
User-agent: *
Disallow: /*?*orderby=
Disallow: /*?*filter_
Disallow: /*?*query_type_
Disallow: /*?*min_price=
Disallow: /*?*max_price=
Disallow: /*?*rating_filter=
Disallow: /wp-content/uploads/wc-logs/
Disallow: /wp-content/uploads/woocommerce_transient_files/
Disallow: /wp-content/uploads/woocommerce_uploads/
Disallow: /*?add-to-cart=
Disallow: /*?*add-to-cart=
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/wp-sitemap.xml
Replace example.com with your site's address. Keep any lines your file already has for other purposes.

What a block does, and what it does not

  • It stops the crawling. Google's crawl budget guide says a robots.txt block prevents Google from crawling an address, and significantly decreases the chance of it being indexed.
  • It does not take an address out of the index. Google's introduction to robots.txt says a blocked address can still be listed if other pages link to it, with no description under it. Search Console has a status for exactly that.
  • Nothing on a blocked page is read. That includes a noindex and a canonical link.
  • It is not instant. Google says it generally keeps a copy of robots.txt for up to 24 hours.

A canonical link answers a different question from a robots.txt rule. A rule decides whether an address is fetched. A canonical link, on a page that was fetched, says which of several similar addresses you would like shown. Google calls it a strong signal for that, and says in its faceted navigation page that it may reduce crawling of the other versions only over time.

The two do not combine on one address, because a blocked page's canonical link is never read. So choose per kind of address. Where the crawling is the problem, block. Where the only question is which address shows in results, a canonical link to the plain address is enough, and so is doing nothing: Google then picks.

If you add canonical links to lists, with a plugin or with code, give each page of a list its own. Google's pagination documentation says not to make the first page the canonical of the pages after it.

How to read the statuses in Search Console

The names below are the ones Google's help for the Page indexing report uses.

StatusWhich of these addresses land thereWhat Google's help saysWhat to do
Alternate page with proper canonical tagA sorted or filtered address whose canonical link, from an SEO plugin, points at the plain oneThe page points correctly to its canonical page, which is indexed, "so there is nothing you need to do"Nothing
Duplicate without user-selected canonicalThe same addresses on a store whose lists print no canonical linkGoogle chose another page as the canonical and will not show this one. It is not an error.Nothing, unless inspecting the address shows Google chose a sorted address over the plain one
Crawled - currently not indexedAny of themCrawled and not indexed. It may or may not be later, and there is no need to resubmit it.Nothing
Discovered - currently not indexedAny of themFound and not yet crawled, typically because Google held back to avoid overloading the siteIf these are a large share of the store's addresses, add the rules above
URL blocked by robots.txtAdd-to-cart addresses, and sorted and filtered ones once the rules are inBlocked by the site's robots.txt, which does not guarantee the address stays out of the indexNothing. This is the rules working.
Indexed, though blocked by robots.txtA blocked address that other pages link toIndexed from the links to it, with a very limited snippetLeave it, or do as the help says to get it out: remove the block and use noindex
URL marked ‘noindex’The cart, checkout and my-account pages, and product searchesGoogle met a noindex and did not index the pageNothing. It is deliberate.
Page with redirect/checkout/, which sends a visitor with an empty cart to /cart/The address redirects, so it will not be indexedNothing
Not found (404)A page number past the end of a list, after products were removedA 404 is not necessarily a problemNothing
Soft 404A filter that matches no product fits the description: it answers 200 and says "No results found"A page that tells the visitor nothing was found, without a 404 statusNothing, or block the filter parameters

What not to do

  • Do not combine noindex with a robots.txt block on the same address. The block stops Google fetching the page, so the noindex is never read. Google's noindex documentation says so directly. Pick one: a block to stop the crawling, or a noindex to keep a crawled page out of results.
  • Do not put noindex in robots.txt. Google's documentation says a noindex rule inside robots.txt is not supported.
  • Do not redirect sorted and filtered addresses to the plain one. A redirect acts on visitors as well as crawlers, so the sort order and the filters stop working for the people using them. Google's canonical documentation keeps redirects for one case, getting rid of a duplicate page for good, and a filtered list is a page you still serve.
  • Do not block the cart, checkout or account pages, for the reason in the warning above.
  • Do not use the removals tool to tidy the report. Google's canonical documentation says not to use URL removal for this, because it hides every version of an address from Search.
  • Do not point page 2 at page 1 with a canonical link. Each page of a list is its own page to Google.

When to get help

  • Ask the host if you cannot reach wp-content or the site's root to add a file.
  • Hand it over if the report grows after the rules are in, or the store's own output does not match this page and you cannot tell which plugin changes it. A WordPress SEO audit from WP Ministry is a written report of what stands between the store and a search engine, each finding with the documentation it rests on, with the fixes on the WordPress side done and tested again.

What a product page itself tells Google is a separate subject, covered in WooCommerce schema markup.

Common questions

Why does Search Console list addresses with ?orderby= when my shop has no such links?

A link is one way Google finds an address, not the only one. Its documentation says a page may be found through other means, such as a link from another site. Once one sorted address is known, the page it returns links to more: the page links on /shop/?orderby=price lead to /shop/page/2/?orderby=price.

Are filter and sort addresses in the report an error?

Not by themselves. Google's help for the report says of "Alternate page with proper canonical tag" that nothing needs doing, and of "Duplicate without user-selected canonical" that it is not an error. It also says a site should not expect every address to be indexed, only the canonical ones.

Should I noindex filtered pages?

It is one of the two ways Google's pagination documentation gives for keeping sorted and filtered lists out of the index. The other is a robots.txt rule, and the two cannot be used on the same address. A noindex does not reduce crawling: Google's crawl budget guide says Google still requests the page and then drops it. WooCommerce has no setting for it on lists, so it takes an SEO plugin or code.

Does WooCommerce block add-to-cart addresses by itself?

In robots.txt, yes, on the version this page was tested on. It adds two lines for add-to-cart to the answer WordPress writes. They are lost when a real robots.txt file exists, because the file replaces that answer. WooCommerce also marks the add to cart control rel="nofollow" wherever it is a link to an add-to-cart address.

Why are my cart and checkout pages marked noindex?

WooCommerce marks them, and the my-account page, on purpose. Its source describes them as dynamic pages that search engines should not index. Leave the tag in place, and leave the pages open to crawlers so that the tag can be read.

I added the rules. Why are the addresses still in the report?

A blocked address does not leave the report. It moves to a blocked status. Google also says it generally keeps a copy of robots.txt for up to 24 hours before it reads the new one, and an address already in the index can stay there under "Indexed, though blocked by robots.txt".

More on this subject

SEO audit, done for you

SEO audit is $249. A written audit, each finding with its source, and up to 2 hours of fixes done and tested again. It starts with a free diagnosis.