A smoke detector doesn't wait for you to smell fire — it sniffs the air constantly and screams the moment something's wrong, while you can still do something about it. Software needs the same early warning.
Health Endpoint Monitoring gives every service a little built-in detector: a special URL that, when pinged, answers one question honestly — am I healthy enough to be handling requests right now? Something on the outside checks it regularly and reacts the instant the answer turns to no.
The problem
Services fail in quiet, sneaky ways. The process is still running, so the operating system thinks all is well — but it's lost its database connection, its disk is full, or a downstream dependency is timing out. From the outside it looks alive while it's actually serving errors.
Without an honest signal of fitness, your traffic router keeps cheerfully sending real users to a sick instance, and you find out it's broken from angry customers rather than from your tools. You need a way to ask each instance, directly and often, whether it should still be in the line of duty.
Step through it below. First the load balancer relies on a shallow check that only proves the process is alive; then flip the switch to a deep check. Each time, predict what the next probe says before you see it.
How it works
Each service publishes a dedicated endpoint — say /health — whose only job is to assess and report fitness. A naïve version just returns 200 OK to prove the process responds. A good one quickly checks the things the service can't work without: can it reach its database, is there free disk, are critical credentials still valid? It rolls that up into a clear verdict: 200 for healthy, 503 for not.
Something then probes that endpoint on a schedule: a load balancer, an orchestrator, or a monitoring service. To avoid flapping on one slow response, it acts only after several failures in a row — three failed probes five seconds apart might take an instance out of rotation, and a couple of passing probes put it back. Pulling an instance should also fire an alert, because a pool that quietly shrinks is an emergency of its own.
Where the probe comes from matters too. A load balancer's probe tells you about each instance. An external monitor that calls your public URL from outside your network, the way a customer would, also catches broken DNS, expired certificates and network faults that no check inside the box can see.
Your load balancer probes /health, which just returns 200. Users report errors from one instance, yet every health check is green. What's the most likely cause?
How deep should the check go?
Deep checks have a trap. Instances usually share their dependencies — the same database, the same downstream APIs. If one of those slows down and every instance's /health checks it, every check fails at the same moment and the load balancer pulls the whole fleet. There's no healthier instance to send traffic to, so pulling them all turns a partial problem into a full outage. (Some load balancers guard against this by failing open: when every instance looks unhealthy, they ignore the checks and keep sending traffic. Don't count on it.)
A useful rule: fail the check only for problems that moving traffic to another instance would fix, and only for things the service truly can't work without. One instance's broken connection pool is a reason to pull it. A slow recommendations API is not: report it as degraded in the response body, so dashboards and alerts see it, and keep answering 200.
Below, every instance's /health also calls a shared Reviews API that only powers a nice-to-have box on product pages. Predict what happens when it slows down, then flip the switch to make Reviews report-only.
A health check can cause the outage it was meant to catch. Put a shared dependency in every instance's must-pass check, and one slow dependency fails them all at once. Probes add load of their own, too: 200 instances, probed every 5 seconds by three load balancer nodes, is 120 checks a second — each one hitting your database if the check queries it.
Make the check meaningful but never expensive. A health endpoint that runs a full query workload or fans out to every dependency on every probe can become a load source of its own — or report sick simply because it timed out under its own weight. Cache dependency checks for a few seconds and put a hard timeout on the whole thing, so it stays a fast, truthful pulse.
When to use it
Health endpoint monitoring is nearly always worth it for anything running in production behind a load balancer or orchestrator — it's the signal those systems rely on to keep traffic flowing to healthy instances. It pairs naturally with a circuit breaker, which stops calling a dependency that keeps failing, and with retry logic that backs off until health is restored.
The main pitfalls are checks that lie: too shallow and they miss real failures, too deep and they pull healthy instances or add load. Tune what "healthy" means to match what the service genuinely needs to do its job, and you get an early-warning system that quietly keeps bad instances away from your users.
Which of these should make an instance's /health answer 503?