The Unwound Geiger Counter: On the Peril of a Quiet Background Check
It's a piece of received wisdom so deeply ingrained in our operational playbooks that we rarely even say it aloud: a background check, ticking away without complaint, is a good thing. We set up our health checks to run on a schedule, a silent sentinel in the code. When they pass, they whisper a quiet 'all is well' into our logs. We aggregate these green ticks into a dashboard, and a sea of green becomes our definition of stability. The system is healthy. We can relax.
But this is where the analogy of a sentinel breaks down, and a more unsettling one takes its place: the Geiger counter. A Geiger counter is an instrument designed not for quietude, but for its opposite. Its purpose is to crackle, to disrupt the silence with evidence of an invisible, potentially dangerous reality. When a Geiger counter is silent, it doesn't mean the environment is safe; it might simply mean the instrument itself is broken, unpowered, or pointed in the wrong direction. A silent Geiger counter in a radioactive field is not a sign of peace, but a catastrophic failure of observation.
Our background health checks suffer from this same peril. We have become so adept at automating the 'everything is okay' signal that we've built a faith-based infrastructure around its silence. We trust that no news is good news. But what if the check is no longer running? What if a subtle code deployment altered its path, or a dependency it relies on has vanished, leaving it to exit gracefully with a 'success' code? The system it was meant to monitor could be on fire, spewing errors to users, while the health check continues to report a serene, and utterly false, positive.
This quiet failure is more dangerous than a loud one. A screaming alert, for all the pain it causes, demands a response. A silent background check, however, breeds complacency. It creates a blind spot that grows in the dark, a sense of security that is entirely unearned. The system drifts into an unknown state, and our primary tool for detecting that drift has fallen asleep at its post, reporting back that the coast is clear even as the tide has already come in.
The solution isn't to abandon background checks, but to mistrust their silence. We must build observability not just for our applications, but for our observability tools themselves. The health check must be monitored. Its successful execution, its latency, and its very existence need to be tracked as fiercely as the services it watches. A missed execution window should trigger an alert just as severe as a 500 error. We need to hear the crackle of the Geiger counter itself, confirming it is alive and listening. For in the complex systems we maintain, a quiet check isn't a sign of health; it's the one signal that should make us most nervous.
Notes & further reading
A few pages I came back to while writing this:
- Washington, DC
- The Bell and the Hammer: On the Tempo of a Health Check
- one area's overview
- The Overwatch Paradox: On the Fallacy of Full Observability
- a practical rundown
- The Lighthouse Keeper's First Mistake: On the Dangers of an Unwritten Log
- Little Rock, AR
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT