Skip to content

WordPress uptime monitoring: what a check sees and what it misses

An uptime monitor is a service outside your server that asks for your site every few minutes and alerts someone when it gets a server error or no answer. Have it send a GET request, look for words on the page, and ask for one address no cache keeps.

By
WP Ministry
Published
Tested on
WordPress 7.1.3, PHP 8.3.35

In short

  • Monitoring has to run somewhere other than your server. A plugin that offers it is a switch for a service that runs elsewhere.
  • WordPress answers 500 for a database error or a critical error, and 503 while updates are installed. A check on the status code sees those.
  • Some broken pages answer 200. A missing theme folder gives an empty page, and an error partway through gives a page that stops. Only a check for words on the page sees those.
  • A check that sends a HEAD request can be told 200 while a visitor gets "There has been a critical error on this website." Set the check to GET.
  • A page cache can answer for a site whose database is down. Add a second check on an address no cache keeps, such as wp-login.php.
  • Send the alert somewhere that does not depend on the server being watched, and decide beforehand who acts on it.

An uptime monitor is a service that runs somewhere other than your server. Every few minutes it asks for an address on your site, and it alerts someone when the answer is a server error or does not come. That much is the same on any website. What is particular to WordPress is what a broken site answers, and two of those answers pass an ordinary check. This page shows each one, and the settings that close the gaps.

What an uptime monitor does

A monitor makes the request a visitor's browser makes, on a schedule, and judges what comes back. It can judge three things:

  • Whether an answer came in time. No connection, a name that cannot be found, an expired certificate and a server that takes too long all fail here.
  • The status code. Every answer begins with a number. 200 means the server says the page is fine. A number from 500 up means the server says it failed.
  • The words on the page. A check can look for a phrase that a healthy page contains, and fail when the phrase is missing.

A check that stops at the status code is the simplest kind. The rest of this page is about why that is not enough on WordPress.

Why monitoring cannot live inside WordPress

WordPress runs on the server being watched. When the server stops answering, nothing inside it can report that. A plugin that offers uptime monitoring is a switch for a service that runs on someone else's servers. Jetpack's documentation says so of its own monitor: once it is switched on, one of Jetpack's servers checks the site every five minutes.

WordPress does send an email about some failures, with the subject "[site name] Your Site is Experiencing a Technical Issue". Its code sends it when a plugin or a theme causes a fatal error on the login page or in the dashboard, and by default no more than once a day. An error that only visitors meet sends nothing. A server that is not answering sends nothing either. And if the site cannot send mail, that email does not leave.

What WordPress answers when it is broken

On a test site running WordPress 7.1.3 on Apache, each of these faults was set up and the home page was asked for with a GET request, as a browser asks.

What is wrongStatusWhat the page says
Updates are being installed503, with Retry-After: 600"Briefly unavailable for scheduled maintenance. Check back in a minute."
WordPress cannot connect to its database500"Error establishing a database connection"
A plugin fails as WordPress starts500"There has been a critical error on this website."
A plugin fails while the page's content is being built500"There has been a critical error on this website."
.htaccess holds a line Apache cannot read500Apache's own page: "Internal Server Error"
A plugin fails after part of the page has been sent200The page stops where the error happened, with no message
The folder of the active theme is missing200Nothing. The page is empty

A check on the status code sees the first five. Each has a page of its own here: scheduled maintenance, the database connection, the critical error and the 500 error.

The last two answer 200, so a check on the status code passes them. A visitor gets a white screen or half a page.

The sixth row has a reason in WordPress's code. Its error handler shows the critical error page, and sets the 500, only when no part of the page has left the server yet. Once the first part of a page has been sent, the status code has gone with it and cannot be changed. On the test site an error raised in the middle of the page's head gave a 200 and a page cut off at that point. That was so with PHP's output buffering off, which is PHP's built-in default, and with it at 4096 bytes, the value in the configuration files PHP ships.

When the server itself is not answering, there is no status code from WordPress at all. The monitor gets a timeout or a refused connection. If a proxy or a CDN stands in front, it answers in the server's place, with a 502 or a 504. Cloudflare's documentation gives the range 520 to 527 for an origin it cannot reach.

See what a monitor sees

You can make the same request from your own computer. This prints the status code and the seconds the answer took. Put your own address in place of example.com.

bash
curl -s -o /dev/null -w "%{http_code} %{time_total}\n" https://example.com/
text
200 0.412345

If the number is 301 or 302, the address you typed redirects to another one. Give a monitor the address at the end of the redirects, with https and with or without www as your site uses it. The redirect checker shows each hop.

This one asks for the headers only. It sends a HEAD request, which matters below.

bash
curl -s -I https://example.com/

On the test site while an update was in progress, its output included these lines:

text
HTTP/1.1 503 Service Unavailable
Retry-After: 600
Cache-Control: no-cache, must-revalidate, max-age=0, no-store, private

And this one counts the lines of the page that hold a phrase. Use words from your own footer.

bash
curl -s https://example.com/ | grep -c "Words from your footer"

It prints 1 or more when the words are there and 0 when they are not. On the test site it printed 0 for every fault in the table.

Six settings that decide what a check can see

Look for each of these in a monitoring service's settings. The names vary from one service to the next.

The address

Start with the home page, at its final address. The second check, further down, is for an address behind it.

The method: GET, not HEAD

A HEAD request asks for the headers of a page without the page. The HTTP standard says a server should answer it with the headers it would send for a GET. WordPress does less than that. On a HEAD request it stops before it builds the page, so an error that would happen while the page is built does not happen.

On the test site, with a plugin failing while the content was built, a GET request got 500 and "There has been a critical error on this website." A HEAD request to the same address, at the same moment, got 200.

Some monitors send HEAD unless told otherwise. Jetpack's documentation says its monitor checks the home page with a HEAD request. UptimeRobot's help says HEAD is its default method, and that GET is a setting some of its plans have. Neither service was run for this page. What was run is the request they describe, and what WordPress answers to it.

Faults that stop WordPress as it starts answer the same to both methods: the database error and the maintenance page gave 500 and 503 to HEAD as well.

Words to look for

Have the check look for a phrase on the page, and fail when it is missing. Choose words from near the end of the page, such as a line of the footer. An empty page lacks them, and so does a page cut short before it reaches them. A page cut after them still has them: on the test site, an error in what WordPress prints after the footer left the footer's words in place. Avoid the site's name: it is in the page's title, at the top, and a page that stops after its head still has it.

A check for words needs the page, so it sends a GET request. That settles the method as well.

The interval

The interval is how long a fault can last before the monitor first sees it. With a check every five minutes, a site can be down for almost five minutes before the first failed check, and an outage shorter than that can pass between two checks unseen.

The timeout and the second look

A single failed check is weak evidence. The request crosses a network that drops things, and a slow answer is different from no answer. A monitor should ask again, or ask from a second place, before it calls an outage. Jetpack's documentation describes this: a failed check marks the site down tentatively, three servers in different locations then check, and an alert goes out only if all three fail.

Where the alert goes

Two questions decide whether an alert is of any use.

  • Does it arrive when the site is down? A mailbox hosted on the same server as the site goes down with it. Send alerts to an address hosted elsewhere, to a phone, or both.
  • Who acts on it, and when? An alert nobody reads until morning is a record. Decide who is told and what they do first. The site down runbook gives the order.

If you use Jetpack's monitor, its documentation says the emails go only to the WordPress.com account that connected Jetpack to the site, and that no other recipient can be added.

A cache can answer for a site that is down

A page cache keeps finished pages and hands them out without running WordPress. That is its job, and it is also how a monitor can be told 200 by a site that is broken.

A page cache plugin works through a file named advanced-cache.php. WordPress loads that file before it connects to the database. On the test site, a small stand-in for such a cache was put in place, holding a copy of the home page, and the database password was then made wrong:

  • the home page answered 200, from the stored copy
  • wp-login.php answered 500, with "Error establishing a database connection"

A cache in front of the server can do the same from further away. Cloudflare's documentation describes a STALE answer, served from its cache when it could not reach the origin server, and a feature named Always Online that shows visitors an archived copy of a page when the origin cannot be reached. Jetpack's documentation says its monitor may not notice an outage behind Cloudflare for this reason.

Maintenance mode is the exception on the test site: WordPress checks for it before it loads the cache file, so the 503 came through. A cache that answers before PHP runs, in the web server or at a CDN, is ahead of that check too.

Add a check on an address no cache keeps

Keep the check on the home page, since that is what visitors get. Add a second one on an address that has to be built every time. wp-login.php is one: it needs PHP and the database, and WordPress sends it with a header that tells caches not to keep it.

bash
curl -s -o /dev/null -w "%{http_code} %{time_total}\n" https://example.com/wp-login.php

Two cautions. If a security plugin has moved the login page, use its new address. And if a firewall limits requests to the login page, the monitor may be answered 403 or 429 there, which says nothing about the site.

On a store, the cart is another such address. WooCommerce's documentation lists Cart, My Account and Checkout as pages to keep out of every cache. A 200 from the cart shows only that the page was built. It does not show that an order can be placed. The checkout down runbook covers that.

To see whether an answer came from a cache, read the headers. Behind Cloudflare, cf-cache-status: HIT and an age line mean the answer came from its cache. Other caches name their own headers.

False alarms, and what causes them

  • Updates. While WordPress installs an update it answers 503, usually for seconds. A check that lands in that moment fails. A second look a minute later clears it. If the 503 stays, WordPress stops honoring the file behind it after ten minutes. Briefly unavailable for scheduled maintenance has the fix.
  • A firewall that turns the monitor away. A monitor is a program asking on a schedule, and a firewall may treat it as a bot. Jetpack's documentation lists a 403 answer as "blocked". Behind Cloudflare, a challenge page carries the header cf-mitigated: challenge, and a monitor cannot pass a challenge. Allow the monitor in the firewall by the name or the addresses its service publishes.
  • A slow answer. A page that takes longer than the monitor's timeout counts as down, though a patient visitor would get it. That is worth knowing in its own right.
  • The monitor's own network. A fault between the monitor and your server looks like your site being down. This is what the second look from another place is for.

What no uptime check can see

  • What happens between two checks. A fault that comes and goes inside the interval may never be seen.
  • What only some visitors meet. A check comes from one place, or a few. A fault in one region, or on one network, can miss it.
  • A page that answers and is wrong. A contact form that sends nothing, a checkout that fails at payment and a page with someone else's links in it all answer 200 with the footer in place. Those need their own checks: the maintenance checklist has the weekly look, and the hacked site check looks at a page from outside for signs of a break-in.

What an uptime percentage is worth

A percentage sounds exact. Work out what it allows in a month of 30 days:

UptimeTime down it allows
99%7 hours 12 minutes
99.9%43 minutes 12 seconds
99.99%4 minutes 19 seconds

A check every five minutes makes 8,640 checks in that month. One failed check is 0.012 percent of them, so a single failure moves the figure from 100 to 99.99. The check cannot tell a fault of ten seconds from one of almost five minutes. Read a percentage together with the interval it was measured at, and ask for the list of outages behind it: when each began and when it ended.

When the alert comes

Look from a second place before you change anything: a phone on mobile data, or the first command on this page. Then follow the site down runbook, which starts with what the screen says and works inward.

An outage also reaches search. Google's documentation says server errors make its crawlers slow down, that pages already in its index are kept at first and eventually dropped, and that crawling picks up again gradually once the server answers normally. The downtime cost calculator works out what an hour costs from your own figures.

If someone else watches the site

A hosting plan or a care plan may include monitoring. Ask the questions this page raises: how often it checks, from where, with which method, whether it reads the page, and who acts on an alert. What WordPress maintenance includes has the full list for each part of a plan.

On our own care plans, monitoring is part of every plan. Our server asks your site's address about every five minutes, and the team is told when two checks in a row get no answer or a server error. A check is made from one place and sees only whether the address answered. It does not read the page, so the two faults that answer 200 would pass it. The WordPress maintenance service page says what else a plan covers.

Common questions

Can a WordPress plugin monitor my site's uptime?

Only as a way to switch on a service that runs elsewhere. Anything that runs on your server stops when the server does, so the checking has to come from outside. A plugin can connect the site to such a service and show its results in the dashboard.

How often should a monitor check?

Every five minutes is common, and every minute is better on a site that takes orders. The interval is the longest a fault can last before the first failed check. Whatever it is, have the monitor look a second time before it alerts.

My monitor says the site is up, but the site is broken. Why?

Four causes are likely. The check sends a HEAD request, and WordPress stops before it builds the page on those. The check reads only the status code, and the broken page answers 200. A cache answered in the site's place. Or the fault is in something the check does not ask for, such as a form or a checkout.

Why did I get an alert while updates were running?

WordPress answers 503 while it installs an update, and a check that lands in those seconds fails. A monitor that looks twice before alerting will not report it. If the alerts last for minutes, the update did not finish.

Should a 403 or a 404 count as down?

For visitors, a 404 on the home page is a broken site, so count it. A 403 given only to the monitor is a firewall turning the monitor away while visitors get the page. Check from your own browser, then allow the monitor in the firewall.

More on this subject

Would you rather we looked after it?

The Essential plan is $39 a month. Updates, backups, monitoring and security scanning. No edit time. It starts with a free diagnosis.