Skip to content

Can this page be indexed? A check from outside

Give a page's address and see what could keep it out of a search engine: the status it answers with, a rule in robots.txt, a noindex on the page or in its headers, and the address it names as its own.

Such as example.com/about/. Our server asks for that page and for the site's robots.txt, as a browser would, and keeps neither the addresses nor what they sent.

What it looks at

Our server asks for the page at the address you give, follows any redirect, and then asks the same site for its robots.txt. It reads four things, each of which can keep a page out of a search engine whatever is written on it.

  • What the address answers with. A page that answers 404, asks for a sign-in or fails with a server error gives a crawler nothing to list. If the address is sent on to another, the page it ends on is the one that counts.
  • What robots.txt says of it. Whether a rule tells Google's or Bing's crawler not to ask for this address, and which line of the file decides. A crawler follows the group of rules that names it and, only when none does, the group for every crawler; the longest rule that matches the address wins.
  • What the page says of itself. A noindex in the page's robots tag, or in a header named X-Robots-Tag that the server sends with it. The header does not show in the page's source, which is why it is so often missed.
  • The address it asks to be listed under. Its canonical link: whether the page names itself, another page of the site, or another site.

What it cannot tell you

It cannot tell you whether the page is in Google. Only Google can: the URL Inspection tool in Search Console says what Google has for an address. This check says whether anything visible from outside stands in the way. When nothing does, whether to list the page is the search engine's own decision, and no setting on a site makes that decision for it.

It also sees the site as our server is shown it, once. A firewall or a host that treats a search engine's crawler differently from other visitors is not something we can see, because the request names itself as ours and never as a crawler.

If the result says the page asks not to be listed and the site runs WordPress, start with the box under Settings, then Reading. For the whole list of checks, in order, see a WordPress site that is not on Google.

What happens to the address

The address goes to our server, which makes the two requests and sends back the report. Neither address, nor what they answered with, is kept or logged. The requests name themselves as WPMinistry-SiteCheck, so they can be told apart in the site's own logs. Only public web addresses are asked.

Common questions

Nothing is in the way, but my page is still not on Google. Why?
Because being allowed in is not the same as being listed. A search engine has to find the page, crawl it and then decide whether to list it. Google's own documentation says crawling can take anywhere from a few days to a few weeks, and that asking for a crawl does not mean a page will be included. Search Console's URL Inspection tool says which stage an address has reached.
robots.txt blocks the page and it has a noindex. Which one wins?
The block, and not in the way people expect. A crawler that is told not to ask for a page never reads it, so it never sees the noindex. Google's documentation says such a page can still appear in results if other pages link to it. To take a page out of search, let it be crawled and keep the noindex.
Why does it check Google's and Bing's crawlers and not others?
Because a robots.txt can give each crawler its own rules, and those two are the ones most sites are asking about. Bing's index is also what Microsoft's Copilot answers from. A rule written for every crawler applies to any crawler that the file does not name.
The page names a different address as its canonical. Is that wrong?
Not always. If that address is the same page without tracking codes or with a tidier form, it is doing its job. If it is a different page, this one is being offered to search engines as a copy of it, and they will usually list the other.
Can I check a page that is not mine?
Yes. It asks for one public page and one public file, which anyone's browser can do, and reads only what they send.

Rather have it found and fixed?

Tell us the address and what this showed. We look at the site itself and reply with the cause and a fixed quote. The diagnosis is free.

Get a free diagnosis