The Watchmaker's Silent Cog: On the Deception of the Passing Health Check
We are taught to worship the green checkmark. In our dashboards, it is the universal symbol of health, a tiny digital god we pray to for peace of mind. A service is up, a check has passed, and all is right with the world. But I want to posit a dangerous, heretical thought: the passing health check is the most successful liar in our observability stack. It tells us a service is alive, but it says nothing of whether it is well. It confirms a heartbeat, but remains silent on the creeping dementia within.
Consider the watchmaker’s finest chronometer. From the outside, its face shows the correct time. Its second hand sweeps with metronomic precision. This is our ‘200 OK’. But inside, a single cog, worn down to a nub, silently fails to engage its neighbor. The mechanism compensates, for now. The time remains correct, the heartbeat steady. But the system is fundamentally degraded, its resilience crumbling. It is only a matter of cycles before the entire assembly jams. Our health check, which merely verifies the second hand is moving, would report perfect health right up until the moment it stops.
This is the great deception. We have optimized for a binary world of up/down, a relic of a simpler time when a crashed process was the primary foe. Today’s failures are subtler. They are the cache that slowly bleeds its entries, leading to rising database load. They is the third-party API that begins returning answers a half-second slower, introducing cascading latency. They are the background thread that silently dies, leaving a queue to grow fat and unprocessed. All the while, the health check—a simple HTTP endpoint returning a happy code—continues to pass with flying colors.
We have been lulled into a false sense of security by our own simplistic definitions of ‘health’. True observability isn’t about verifying a pulse; it’s about conducting a full medical exam. It requires interrogating the state beyond the binary. It demands we watch the length of the queue, the latency of the dependency, the rate of the errors, the composition of the traffic. A service can be ‘up’ but drowning, and a passing check is the smile it gives us as it sinks beneath the waves.
The received wisdom we must critique is that a health check is a goal unto itself. It is not. It is the barest minimum, the first and most primitive of many signals. To build truly reliable services, we must learn to listen for the silence—for what the passing check does *not* say. We must become diagnosticians, not mere pulse-takers, and build systems that can tell us not just that they are running, but how they are living.
Notes & further reading
A few pages I came back to while writing this: