The Gardener's Cursed Pruning Shears: On the Folly of Unquestioned Health Checks
We're told to tend to our services like a gardener tends to a prized rose bush. We prune the dead wood, check for blight, and ensure the roots are healthy. In our world, the most fundamental tool for this is the health check: a simple, automated probe that asks, "Are you alive?" It's received wisdom, an article of faith. If the endpoint returns a 200, the service is healthy. If it times out or returns a 500, it is not. It seems perfectly logical, a pristine pair of shears for keeping our digital gardens in order.
But I've come to view this simplicity as a kind of curse. We've become so reliant on the binary green/red light that we've forgotten what health truly means in a complex, interdependent ecosystem. A service can return a perfect 200 status code while being critically ill. It can faithfully respond to your synthetic ping from a health-check container on the same cluster, all while its connection pool is leaking, its internal queues are backing up, and its ability to talk to the user database has silently failed. The shears make a clean cut, declaring the branch healthy, even as the sap within has stopped flowing.
This creates a dangerous illusion of observability. We congratulate ourselves on our 99.99% uptime, a metric derived entirely from these superficial pings, while real users are encountering slow page loads, failed transactions, and cryptic errors. The health check has become a performance, a well-rehearsed piece of theater that the service acts out for its overseers. It’s the equivalent of a patient who can still smile and say "I’m fine" to a doctor’s quick check, while a deeper, systemic illness goes undiagnosed.
The deeper folly lies in how these checks shape our architecture. We design systems to pass the test, not necessarily to be robust. We add trivial endpoints that do nothing but check if the process is running, avoiding any meaningful validation of downstream dependencies for fear of causing a cascade of failures. In doing so, we optimise for the metric of "availability" at the expense of the actual experience of "reliability." We're pruning the bush to look good from the garden gate, without checking if the flowers have any scent.
So, what's the alternative? It's not to abandon health checks, but to mistrust them profoundly. We must see them not as a definitive diagnosis, but as the first, most basic clue. True health is a multidimensional state. It requires the corroborating evidence of business logic metrics—like successful login rates or checkout completions. It demands the rhythm of real-user monitoring, showing the actual latency and error rates experienced by people, not just bots. It needs the deep tissue scan provided by distributed tracing, following a request through the tangled root system of microservices. Our shears are just one tool; we need the entire shed.
The cursed pruning shears promise a clean, simple answer. But a living system is never simple. By fetishizing the green light of a 200 response, we risk cultivating a garden of beautiful, hollow plants that look impeccable until the moment they collapse. True reliability is messier, harder to measure, and requires a gardener who is willing to get their hands dirty, looking beyond the easy cut to understand the true vitality of the system.
Notes & further reading
A few pages I came back to while writing this: