---
title: Health checks and monitoring
summary: A green status page and a broken app can happen at the same time. What a health check can see, what it cannot, and how to count the failures real users hit.
topic: infrastructure
order: 3
updated: 2026-09-19
---

## What it is

A **health check** is a small URL a monitor calls every minute or so to ask "are you alive?".
**Monitoring** is everything you watch to know whether the service works for the people using
it: errors, response times, memory, the database.

## Why it is a rule

- **Amazon S3, 28 February 2017.** An engineer ran a command to remove a few servers, and a
  typo removed far more. S3 in its biggest region went down for about four hours and took a
  large part of the internet with it. AWS's own status page kept showing green for a while,
  because the status page itself depended on S3.

Two lessons: the monitor must not depend on the thing it monitors, and "the front door
answers" is not the same as "the service works".

## How to do it

### Separate liveness from readiness

```text
GET /health      → the process is up and answering          (liveness)
GET /health/db   → the process can reach the database        (readiness)
```

A platform restarts a process that fails **liveness**. It stops sending traffic to one that
fails **readiness**, but leaves it running, because restarting will not fix a database that
is down.

Keep the liveness check cheap. A health check that runs a heavy query becomes the load that
takes the server down.

### Probe from outside

The monitor runs somewhere else: another provider, another region. A probe inside the same
server cannot tell you that the server is unreachable.

### Count errors where they happen

A probe only sees the URLs it asks for. Real users hit hundreds of routes the probe never
touches, so a module can fail all morning while every check stays green.

```js
function tryCatch(fn) {
  return (req, res, next) =>
    Promise.resolve(fn(req, res, next)).catch((err) => {
      noteError(req, err)
      res.status(500).json({ error: 'Something went wrong' })
    })
}
```

If every route goes through one wrapper, that wrapper is the one place that sees every
failure. Count them there.

### Report a rolling window, not a total

"37 errors since boot" means nothing. "12 errors in the last 15 minutes, all in billing"
says what is broken and that it is broken now.

### Keep the health endpoint quiet

It should not reveal versions, paths or stack traces to strangers. Protect detailed health
data with a key, and send `no-store` so no cache keeps an old answer.

## How we do it here

The app exposes a cheap liveness check and a separate database check. They are probed from a
separate Cloudflare Worker, outside the hosting provider. Every route's errors are counted in
one wrapper and reported as a 15-minute rolling window per app.

## Benefits

- An outage is noticed by a machine in a minute, not by a client in an hour.
- Error counts per app point at which part is broken, before anyone opens a log.
- Readiness checks stop traffic going to a half-working server.

## Disadvantages

- Probes cost requests, and an external monitor is one more service to pay for and keep
  running.
- Too many alerts and people learn to ignore them. Alert on what needs a human now, and log
  the rest.
- A health check is a public URL; left unguarded it can leak internals or be used for load.

## Checklist

- Separate liveness and readiness checks.
- Probed from outside the hosting provider.
- Errors counted in one place, per app, over a rolling window.
- Health responses are `no-store` and reveal nothing sensitive.

## Sources

- [AWS — Summary of the Amazon S3 Service Disruption, 28 February 2017](https://aws.amazon.com/message/41926/)
- [Google SRE Book — Monitoring Distributed Systems](https://sre.google/sre-book/monitoring-distributed-systems/)
