WooCommerce filter URLs indexed in Google: ?orderby, ?filter_ and add-to-cart addresses, and what to do
A WooCommerce store answers at extra addresses for one list of products, such as ?orderby=, ?filter_, ?min_price=, /page/2/ and ?add-to-cart=. Each returns the page with no canonical link and no noindex. A small store can leave them. To stop the crawling, add robots.txt rules for the parameters.
- By
- WP Ministry
- Published
- Tested on
- WordPress 7.1.3, WooCommerce 11.2.0, PHP 8.3.35
In short
- A sorted or filtered shop address answers 200 with the plain page's robots tag and no canonical link. WordPress prints a canonical link on single pages only, and WooCommerce adds none to a list.
- WooCommerce already asks crawlers to stay away from add-to-cart addresses, in the robots.txt that WordPress answers with. A real robots.txt file drops those lines.
- The cart, checkout and my-account pages carry noindex on purpose. Never block them in robots.txt.
- For filter addresses nobody needs in search, Google's first answer is a robots.txt rule. It calls a canonical link and rel="nofollow" less effective in the long term.
- Most of the statuses these addresses sit under in Search Console are not errors and need nothing done.
- Do not put noindex on an address that robots.txt blocks, and do not redirect filter addresses to the plain one.
The addresses with question marks in your page indexing report are your shop and category pages, sorted or filtered. WooCommerce makes them from parameters: orderby for the sort order, names that begin filter_ and query_type_ for attribute filters, min_price and max_price, rating_filter, and add-to-cart. Each one answers with status 200 and the same products in another order, or fewer of them. On a store with no SEO plugin, none of them carries a canonical link or a noindex.
For most stores the report is describing something that works as intended, and nothing needs fixing. When the addresses run to thousands, or crawlers spend the server's time on them, the remedy Google names first is a robots.txt rule that keeps crawlers out. This page lists the addresses, shows what each one tells a search engine, and gives the rules.
The addresses one list of products answers at
These were read from a test store with twenty products, one category named Mugs and one attribute named Color, on WordPress's default theme. The plain addresses are /shop/ and /product-category/mugs/.
| Parameter | Example | What it does | What makes a link to it |
|---|---|---|---|
orderby | /shop/?orderby=price | Sorts the list. The values on offer are popularity, rating, date, price and price-desc. | The sort drop-down, which is a form and not a link |
filter_<attribute> with query_type_<attribute> | /shop/?filter_color=red&query_type_color=or | Keeps the products with that value. query_type_ says whether several values mean any of them or all of them. | An attribute filter |
min_price, max_price | /shop/?min_price=5&max_price=10 | Keeps the products in a price range | A price filter, which is a form or a slider |
rating_filter | /shop/?rating_filter=5 | Keeps the products with that average rating | A rating filter |
filter_stock_status | /shop/?filter_stock_status=instock | Keeps the products with that stock status | The Product Filters block |
categories, tags, brands | /shop/?categories=mugs | Keeps the products in a category, a tag or a brand | The Product Filters block |
| A page number in the path | /shop/page/2/ | Shows the next page of the list | The page links under the list |
add-to-cart | /shop/?add-to-cart=10 | Puts the product with ID 10 in a cart, then shows the page | The add to cart control |
Every one of them works on a category page too: /product-category/mugs/?orderby=price is the category, sorted. A product's own page has a further family, the attribute_ parameters that choose a variation, and WooCommerce variations and search covers those.
Which of them a fresh store prints as links
Almost none. On the default theme, the only links on the shop page that carry a parameter or a page number are the page links. The sort control is a drop-down inside a form. The add to cart control is a button with no address in it. There is no filter on the page until someone adds one.
That matters because of how Google finds addresses. Its documentation on pagination says it generally crawls the addresses found in the href of <a> elements, and that its crawlers do not click buttons.
Filters come in two kinds, and only the older kind prints links:
- The classic widgets. "Filter Products by Attribute" prints each value as a link, and marks each link
rel="nofollow". "Filter Products by Rating" prints each rating as a link with norelat all. "Filter Products by Price" is a form. "Active Product Filters" prints a link for removing each filter, markedrel="nofollow". - The Product Filters block. It prints checkboxes and buttons, and changes the address with a script when one is pressed. Its choices are not links in the page.
The add to cart control becomes a link in two cases: a theme or a shortcode that uses WooCommerce's classic product list, and a store where "Enable AJAX add to cart buttons on archives" is unticked under WooCommerce, then Settings, then Products. In both, the link is ?add-to-cart= and the product's ID, marked rel="nofollow".
One parameter leads to more
WordPress's page links keep whatever follows the question mark. On /shop/?orderby=price-desc, the link to the second page is /shop/page/2/?orderby=price-desc, and it carries no rel="nofollow".
That holds for a parameter WooCommerce has never heard of. /shop/?anything=1 answers 200 with the full list, and links to /shop/page/2/?anything=1. So a crawler that meets one such address, on your site or in a link from someone else's, is handed more of them. Google's documentation on link attributes makes the same point from the other side: a page behind a nofollow link may still be found through other means, such as a link from another site.
What each address tells a search engine
This is what WooCommerce and WordPress print by themselves, with no SEO plugin, as read from the test store.
| Address | Status | Canonical link | Robots tag |
|---|---|---|---|
/shop/ | 200 | None | max-image-preview:large |
/shop/?orderby=price | 200 | None | max-image-preview:large |
/shop/?filter_color=red | 200 | None | max-image-preview:large |
/shop/?min_price=5&max_price=10 | 200 | None | max-image-preview:large |
/shop/?rating_filter=4, which matches no product | 200, with "No results found" | None | max-image-preview:large |
/shop/page/2/ | 200 | None | max-image-preview:large |
/shop/page/3/, past the last page | 404 | None | max-image-preview:large |
/product-category/mugs/?orderby=price | 200 | None | max-image-preview:large |
/shop/?add-to-cart=10 | 200, and a cart is made | None | max-image-preview:large |
/product/mug-1/?add-to-cart=10 | 200, and a cart is made | The product's plain address | max-image-preview:large |
/?s=mug&post_type=product, a product search | 200 | None | noindex, follow, max-image-preview:large |
/cart/ | 200 | Itself | max-image-preview:large, noindex, follow |
/checkout/ with an empty cart | 302, to /cart/ | ||
/checkout/ with a product in the cart | 200 | Itself | max-image-preview:large, noindex, follow |
/my-account/ | 200 | Itself | max-image-preview:large, noindex, follow |
Four things to take from it:
- The robots tag blocks nothing.
max-image-preview:largeis on every WordPress page and allows large image previews. The word that keeps a page out of search isnoindex, and no sorted or filtered address has it. WordPress itself puts it on search results, which is why a product search does. - A list has no canonical link. WordPress prints one for a single post, page or product: its
rel_canonical()function stops at once for anything else. A shop page and a category page are lists, so neither gets one, with a parameter or without, and WooCommerce adds none. A product does, which is why an add-to-cart address on a product page points back at the plain product. - A filter that matches nothing is still a page. It answers 200 and says "No results found". A page number past the end answers 404.
- An add-to-cart address is not a view of a page. Each request makes a cart: WooCommerce sets its cart cookies and saves a session in the database. It is the one address here that changes something on the server every time a crawler asks for it.
A theme, an SEO plugin or a filter plugin changes this table, so read your own store's. This prints the canonical link and the robots tag of one address, as a crawler receives them:
curl -s "https://example.com/shop/?orderby=price" | grep -i -o -E '<(link|meta)[^>]*(canonical|robots)[^>]*>'On the test store it prints one line, the robots tag, and no canonical link:
<meta name='robots' content='max-image-preview:large' />An SEO plugin can fill the canonical column. Yoast SEO's published specification, for one, says a canonical link belongs on list pages as well, with sorting and filtering queries taken off it and the page number kept. On a store like that, a sorted address points at the plain one.
The cart, checkout and account pages
WooCommerce puts noindex on the cart, checkout and my-account pages on purpose: its function wc_page_no_robots adds the word through WordPress's own robots tag, and the table shows it on all three. When Search Console lists them under a noindex status, the store is working as designed and there is nothing to fix.
What Google says about filter addresses
Google documents this under the name faceted navigation. Its account of the problem is that filters built on URL parameters can make a near-endless number of addresses. A crawler cannot tell whether one is useful without fetching it, so it fetches a very large number of them before it works out that they are not. That leaves less time for the new addresses that are useful.
It then sorts sites into two. A site that does not need its filter addresses in search should prevent them being crawled. A site that does need them should follow a short list of practices, and accept the cost in server resources.
To prevent crawling, it gives these ways, in this order:
| Google's way | What Google says of it | On a WooCommerce store |
|---|---|---|
| robots.txt rules that disallow the parameters | There is often no good reason to let filtered lists be crawled. Allow the items' own pages and one listing page with no filter. | WooCommerce writes rules for add-to-cart and for nothing else. The rest are yours to add. |
URL fragments for filters, the part of an address after # | Google Search generally does not use fragments in crawling or indexing, so filters kept there have no effect on crawling. | WooCommerce's filters use parameters. No setting moves them to fragments. |
| A canonical link to the unfiltered address | May, over time, decrease how much the other versions are crawled. Generally less effective in the long term than the two above. | No list page has a canonical link unless a plugin adds one. |
rel="nofollow" on links to filtered pages | May help, but every link to an address has to carry it. Also generally less effective in the long term. | Only some links carry it: the attribute widget's and the add to cart link do, and the rating widget's and the page links do not. |
For a site that wants its filter addresses in search, Google asks that a filter combination with no results answer 404, and a page number that does not exist too. WooCommerce does the second and not the first, as the table above shows.
What you can do
Leave it alone when the numbers are small
Google's guide to crawl budget says who it is written for: sites with more than a million pages that change weekly, sites with more than ten thousand that change daily, and sites where a large share of all addresses sit under "Discovered - currently not indexed". It calls those figures rough estimates. It also says that if your pages are crawled the day they are published, you do not need the guide.
Its documentation on canonical addresses says a site will likely do fine without stating a preference, because Google then chooses the version to show. Its help for the Page indexing report says not to expect every address on a site to be indexed, only the canonical ones.
So a store with a few hundred products, whose report lists a few hundred sorted and filtered addresses under the statuses in the table further down, is in the state Google describes as normal. Act when one of these is true: a large share of the store's addresses are "Discovered - currently not indexed", new products wait a long time to be crawled, or the server's logs show crawlers spending their visits on filter addresses.
See what your robots.txt already says
curl -s https://example.com/robots.txtOn the test store, with nothing added, the answer is this:
User-agent: *
Disallow: /wp-content/uploads/wc-logs/
Disallow: /wp-content/uploads/woocommerce_transient_files/
Disallow: /wp-content/uploads/woocommerce_uploads/
Disallow: /*?add-to-cart=
Disallow: /*?*add-to-cart=
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xmlThe five lines after User-agent: * are WooCommerce's. It adds them to the answer WordPress writes when the site has no robots.txt file, and two of them ask crawlers to leave every add-to-cart address alone. Nothing in this answer mentions orderby or a filter.
If your store's answer has no add-to-cart lines, a real robots.txt file or a plugin's editor has replaced WordPress's answer, and WooCommerce's lines went with it. Where robots.txt is in WordPress explains which of the two you have.
Add rules for the sort and filter parameters
A filter on WordPress's robots.txt answer keeps WordPress's lines and WooCommerce's, and adds yours. The smallest place for it is a must-use plugin: one PHP file in wp-content/mu-plugins/, which WordPress loads on every request with nothing to activate.
Step 1: Create the mu-plugins folder if it is not there
Look in
wp-content. If there is no folder namedmu-plugins, create one.Step 2: Add the file
Save this as
store-robots-rules.phpin that folder. It puts six lines into the group that WordPress and WooCommerce already wrote.wp-content/mu-plugins/store-robots-rules.php<?php /** * Plugin Name: Store robots.txt rules * Description: Asks crawlers to stay out of the sorted and filtered copies of the shop's pages. */ add_filter( 'robots_txt', function ( $output ) { $rules = array( 'Disallow: /*?*orderby=', 'Disallow: /*?*filter_', 'Disallow: /*?*query_type_', 'Disallow: /*?*min_price=', 'Disallow: /*?*max_price=', 'Disallow: /*?*rating_filter=', ); return preg_replace( '/^User-agent: \*\s*$/m', "User-agent: *\n" . implode( "\n", $rules ), $output, 1 ); }, 20 );Step 3: Load /robots.txt
Run the command from the last section again, or open the address in a browser. The six lines sit directly under
User-agent: *, above WooCommerce's.textUser-agent: * Disallow: /*?*orderby= Disallow: /*?*filter_ Disallow: /*?*query_type_ Disallow: /*?*min_price= Disallow: /*?*max_price= Disallow: /*?*rating_filter= Disallow: /wp-content/uploads/wc-logs/ Disallow: /wp-content/uploads/woocommerce_transient_files/ Disallow: /wp-content/uploads/woocommerce_uploads/ Disallow: /*?add-to-cart= Disallow: /*?*add-to-cart= Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/wp-sitemap.xmlStep 4: Test the rules against real addresses
Copy the whole answer into the robots.txt tester, which works in your browser. Give it
/shop/, a product's address and/shop/page/2/: each should come back allowed. Then give it a sorted address and a filtered one from your report: each should come back blocked, with the line that decides.
In Google's reading of robots.txt, * stands for any run of characters. /*?*orderby= therefore matches any address that has orderby= somewhere after its question mark, whatever comes before it. Google's own example for faceted navigation is written the same way. filter_ and query_type_ have no = because the attribute's name follows them, and filter_ also covers filter_stock_status.
The rules leave three of the block's parameters out: categories, tags and brands. Those words are common enough in other plugins' addresses that a rule for them should be a decision. Add a line in the same form if your report shows them.
The file is listed on the Plugins screen under "Must-Use" and cannot be switched off there. To undo the change, delete the file.
If the site has a real robots.txt file
The filter does nothing then, because WordPress is never asked. Put the rules in the file itself, with the lines WordPress and WooCommerce would have written, since the file is the whole answer.
User-agent: *
Disallow: /*?*orderby=
Disallow: /*?*filter_
Disallow: /*?*query_type_
Disallow: /*?*min_price=
Disallow: /*?*max_price=
Disallow: /*?*rating_filter=
Disallow: /wp-content/uploads/wc-logs/
Disallow: /wp-content/uploads/woocommerce_transient_files/
Disallow: /wp-content/uploads/woocommerce_uploads/
Disallow: /*?add-to-cart=
Disallow: /*?*add-to-cart=
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xmlWhat a block does, and what it does not
- It stops the crawling. Google's crawl budget guide says a robots.txt block prevents Google from crawling an address, and significantly decreases the chance of it being indexed.
- It does not take an address out of the index. Google's introduction to robots.txt says a blocked address can still be listed if other pages link to it, with no description under it. Search Console has a status for exactly that.
- Nothing on a blocked page is read. That includes a
noindexand a canonical link. - It is not instant. Google says it generally keeps a copy of robots.txt for up to 24 hours.
What a canonical link does here
A canonical link answers a different question from a robots.txt rule. A rule decides whether an address is fetched. A canonical link, on a page that was fetched, says which of several similar addresses you would like shown. Google calls it a strong signal for that, and says in its faceted navigation page that it may reduce crawling of the other versions only over time.
The two do not combine on one address, because a blocked page's canonical link is never read. So choose per kind of address. Where the crawling is the problem, block. Where the only question is which address shows in results, a canonical link to the plain address is enough, and so is doing nothing: Google then picks.
If you add canonical links to lists, with a plugin or with code, give each page of a list its own. Google's pagination documentation says not to make the first page the canonical of the pages after it.
How to read the statuses in Search Console
The names below are the ones Google's help for the Page indexing report uses.
| Status | Which of these addresses land there | What Google's help says | What to do |
|---|---|---|---|
| Alternate page with proper canonical tag | A sorted or filtered address whose canonical link, from an SEO plugin, points at the plain one | The page points correctly to its canonical page, which is indexed, "so there is nothing you need to do" | Nothing |
| Duplicate without user-selected canonical | The same addresses on a store whose lists print no canonical link | Google chose another page as the canonical and will not show this one. It is not an error. | Nothing, unless inspecting the address shows Google chose a sorted address over the plain one |
| Crawled - currently not indexed | Any of them | Crawled and not indexed. It may or may not be later, and there is no need to resubmit it. | Nothing |
| Discovered - currently not indexed | Any of them | Found and not yet crawled, typically because Google held back to avoid overloading the site | If these are a large share of the store's addresses, add the rules above |
| URL blocked by robots.txt | Add-to-cart addresses, and sorted and filtered ones once the rules are in | Blocked by the site's robots.txt, which does not guarantee the address stays out of the index | Nothing. This is the rules working. |
| Indexed, though blocked by robots.txt | A blocked address that other pages link to | Indexed from the links to it, with a very limited snippet | Leave it, or do as the help says to get it out: remove the block and use noindex |
| URL marked ‘noindex’ | The cart, checkout and my-account pages, and product searches | Google met a noindex and did not index the page | Nothing. It is deliberate. |
| Page with redirect | /checkout/, which sends a visitor with an empty cart to /cart/ | The address redirects, so it will not be indexed | Nothing |
| Not found (404) | A page number past the end of a list, after products were removed | A 404 is not necessarily a problem | Nothing |
| Soft 404 | A filter that matches no product fits the description: it answers 200 and says "No results found" | A page that tells the visitor nothing was found, without a 404 status | Nothing, or block the filter parameters |
What not to do
- Do not combine
noindexwith a robots.txt block on the same address. The block stops Google fetching the page, so thenoindexis never read. Google's noindex documentation says so directly. Pick one: a block to stop the crawling, or anoindexto keep a crawled page out of results. - Do not put
noindexin robots.txt. Google's documentation says a noindex rule inside robots.txt is not supported. - Do not redirect sorted and filtered addresses to the plain one. A redirect acts on visitors as well as crawlers, so the sort order and the filters stop working for the people using them. Google's canonical documentation keeps redirects for one case, getting rid of a duplicate page for good, and a filtered list is a page you still serve.
- Do not block the cart, checkout or account pages, for the reason in the warning above.
- Do not use the removals tool to tidy the report. Google's canonical documentation says not to use URL removal for this, because it hides every version of an address from Search.
- Do not point page 2 at page 1 with a canonical link. Each page of a list is its own page to Google.
When to get help
- Ask the host if you cannot reach
wp-contentor the site's root to add a file. - Hand it over if the report grows after the rules are in, or the store's own output does not match this page and you cannot tell which plugin changes it. A WordPress SEO audit from WP Ministry is a written report of what stands between the store and a search engine, each finding with the documentation it rests on, with the fixes on the WordPress side done and tested again.
What a product page itself tells Google is a separate subject, covered in WooCommerce schema markup.
Common questions
Why does Search Console list addresses with ?orderby= when my shop has no such links?
A link is one way Google finds an address, not the only one. Its documentation says a page may be found through other means, such as a link from another site. Once one sorted address is known, the page it returns links to more: the page links on /shop/?orderby=price lead to /shop/page/2/?orderby=price.
Are filter and sort addresses in the report an error?
Not by themselves. Google's help for the report says of "Alternate page with proper canonical tag" that nothing needs doing, and of "Duplicate without user-selected canonical" that it is not an error. It also says a site should not expect every address to be indexed, only the canonical ones.
Should I noindex filtered pages?
It is one of the two ways Google's pagination documentation gives for keeping sorted and filtered lists out of the index. The other is a robots.txt rule, and the two cannot be used on the same address. A noindex does not reduce crawling: Google's crawl budget guide says Google still requests the page and then drops it. WooCommerce has no setting for it on lists, so it takes an SEO plugin or code.
Does WooCommerce block add-to-cart addresses by itself?
In robots.txt, yes, on the version this page was tested on. It adds two lines for add-to-cart to the answer WordPress writes. They are lost when a real robots.txt file exists, because the file replaces that answer. WooCommerce also marks the add to cart control rel="nofollow" wherever it is a link to an add-to-cart address.
Why are my cart and checkout pages marked noindex?
WooCommerce marks them, and the my-account page, on purpose. Its source describes them as dynamic pages that search engines should not index. Leave the tag in place, and leave the pages open to crawlers so that the tag can be read.
I added the rules. Why are the addresses still in the report?
A blocked address does not leave the report. It moves to a blocked status. Google also says it generally keeps a copy of robots.txt for up to 24 hours before it reads the new one, and an address already in the index can stay there under "Indexed, though blocked by robots.txt".
- Error fixHow to fix WooCommerce payment gateway errors
- ResourceWooCommerce checkout is down: a runbook
- Cost guideWooCommerce maintenance cost: why a store costs more, and what the extra should buy
- ResourceWooCommerce sale readiness checklist: Black Friday or any big sale
- Error fixHow to fix a WooCommerce checkout that is not working
- GuideWooCommerce variations and SEO: variation URLs, canonical links and the price in structured data

