The Unkindness of Constant Vigilance: How Health Checks Can Make Your Service Less Healthy
We treat uptime monitoring and health checks as an unalloyed good. The logic is seductive: to know your service, you must poll it. To ensure it's alive, you must ask. We build intricate networks of synthetic probes, each a tiny, well-meaning question: "Are you there? Are you well?" But what if this constant interrogation is itself a source of ill health? What if, in our quest for reliability, we are adding a subtle, corrosive load that erodes the very resilience we seek?
Consider the scale. A global service might have dozens of monitoring nodes, each firing checks every thirty seconds—some every five. Each check is trivial: a database ping, a cache query, a small API call. Individually, they are nothing. Collectively, they are a persistent drizzle, an artificial traffic pattern that never sleeps. In a state of nominal health, this is background noise. But services are rarely in a state of nominal health. They exist in a dynamic tension, scaling up and down, encountering garbage collection pauses, experiencing regional blips, rebalancing data. It is precisely in these moments of delicate convalescence—when the service is recovering from a spike, or a pod is warming up, or a leader is failing over—that our barrage of friendly questions arrives with the grace of a door-to-door salesman during a power outage.
The Check That Kills the Patient
The perverse outcome is this: a health check can trigger a cascading failure it was designed to prevent. A service under duress, already struggling with real user load, must now also satisfy the synthetic, impatient demands of its monitors. A single, poorly-considered check that does a full authentication cycle or runs a complex query can be the straw that breaks the camel's back. The monitor sees a timeout, declares an outage, and triggers an automated remediation—perhaps restarting the struggling instance—which only adds more chaos to the system. Our sentinel has become an active participant in the collapse.
This leads to a counterintuitive, almost heretical principle: the most reliable health check is sometimes the one you don't run. Or, more accurately, the one you run with profound empathy for the system's state. We must design checks that are not just technically correct, but contextually kind. A check should be the lightest possible touch—a glance, not a handshake. It should be intelligent enough to back off when it senses distress, not double down. It should, in its own small way, be part of the cure, not part of the disease.
Observability is about understanding the truth of your system. The naive truth is that constant checking equals constant knowledge. The deeper, more uncomfortable truth is that our instruments interact with the observed system. We are not passive watchers from a soundproof booth; we are in the room, breathing the same air. To run a truly reliable service, we must first ensure our watchers are not a source of its unreliability. We must have the courage, at times, to look away, to trust the system enough to let it breathe on its own. The kindest vigilance knows when to blink.
Notes & further reading
A few pages I came back to while writing this:
- New Orleans, LA
- The Night Watchman's Lantern: What the Rounds Teach Us About Coverage
- Shreveport, LA
- The Unseen Ink: On the Integrity of a Single, Quiet Check
- Boston, MA
- The Weaver's Patience: On Understanding Your Service's True Rhythm
- Springfield, MA
- Worcester, MA
- Baltimore, MD
- Detroit, MI
- Grand Rapids, MI
- Sterling Heights, MI
- Warren, MI