Can this page be indexed? A check from outside
Give a page's address and see what could keep it out of a search engine: the status it answers with, a rule in robots.txt, a noindex on the page or in its headers, and the address it names as its own.
What it looks at
Our server asks for the page at the address you give, follows any redirect, and then asks the same site for its robots.txt. It reads four things, each of which can keep a page out of a search engine whatever is written on it.
- What the address answers with. A page that answers 404, asks for a sign-in or fails with a server error gives a crawler nothing to list. If the address is sent on to another, the page it ends on is the one that counts.
- What robots.txt says of it. Whether a rule tells Google's or Bing's crawler not to ask for this address, and which line of the file decides. A crawler follows the group of rules that names it and, only when none does, the group for every crawler; the longest rule that matches the address wins.
- What the page says of itself. A
noindexin the page's robots tag, or in a header named X-Robots-Tag that the server sends with it. The header does not show in the page's source, which is why it is so often missed. - The address it asks to be listed under. Its canonical link: whether the page names itself, another page of the site, or another site.
What it cannot tell you
It cannot tell you whether the page is in Google. Only Google can: the URL Inspection tool in Search Console says what Google has for an address. This check says whether anything visible from outside stands in the way. When nothing does, whether to list the page is the search engine's own decision, and no setting on a site makes that decision for it.
It also sees the site as our server is shown it, once. A firewall or a host that treats a search engine's crawler differently from other visitors is not something we can see, because the request names itself as ours and never as a crawler.
If the result says the page asks not to be listed and the site runs WordPress, start with the box under Settings, then Reading. For the whole list of checks, in order, see a WordPress site that is not on Google.
What happens to the address
The address goes to our server, which makes the two requests and sends back the report. Neither address, nor what they answered with, is kept or logged. The requests name themselves as WPMinistry-SiteCheck, so they can be told apart in the site's own logs. Only public web addresses are asked.
Common questions
Nothing is in the way, but my page is still not on Google. Why?
robots.txt blocks the page and it has a noindex. Which one wins?
Why does it check Google's and Bing's crawlers and not others?
The page names a different address as its canonical. Is that wrong?
Can I check a page that is not mine?
Rather have it found and fixed?
Tell us the address and what this showed. We look at the site itself and reply with the cause and a fixed quote. The diagnosis is free.
Get a free diagnosis
