The Watchmaker's Empty Bench: On the Obsession with Mean Time
There’s a chart that rules the world of service reliability. It’s the one showing our uptime percentage, that steady line creeping asymptotically towards the coveted 100%. We fetishize this number, we promise it to users, we build our professional reputations upon it. We have become obsessive horologists, polishing the glass and winding the spring, convinced that a perfect mean time between failures is the ultimate achievement. But I’ve come to believe this single-minded focus is like a master watchmaker who spends all their time calibrating the central gear while the countless smaller wheels, the very things that give the hands their purpose, slowly seize up from neglect.
The received wisdom is simple: high availability is the goal, and the Mean Time Between Failures (MTBF) is its primary measure. We pour immense effort into maximizing this number, celebrating ever-longer stretches of uninterrupted service. But this obsession creates a dangerous illusion of health. A service can have a fantastic MTBF while being riddled with what I’ve started to call "silent decays"—performance regressions that shrink our response-time budget by milliseconds each week, or third-party API integrations that degrade so gradually no single event trips an alert. The system remains "up" in the most technical, brutalist sense, yet its utility and user satisfaction are quietly eroding.
This happens because our focus on mean time between catastrophic failures ignores the mean time of everything else. It ignores the mean time to detect a subtle anomaly, not just an outage. It ignores the mean time for a user to complete a critical journey through our application. A service can be technically available but so sluggish that it’s functionally broken for the people who matter most. By concentrating on the big, obvious stoppage, we risk tuning our monitoring to be deaf to the chorus of whispers that precede a shout.
True reliability isn’t just the absence of a siren; it’s the presence of a predictable, quality experience. It’s the texture of the service, not just its binary state. An older mechanical clock, with a slightly irregular tick-tock, might still tell the time perfectly well over a long period, its character defined by the gentle, audible rhythm of its workings. Our modern digital clocks, in their silent, precise pursuit of perfection, offer no such reassurance. They are either correct or they are blank.
Perhaps we need to spend less time at the empty bench, admiring the single, averaged number, and more time listening to the intricate symphony of the gears. We should measure not just the time between failures, but the quality of the time in between. This means embracing a richer set of observability signals—not just "is it up?" but "how is it feeling?" Latency distributions, error rates by user cohort, business transaction success rates—these are the rhythms of a living system. They tell us if the watch is merely keeping time, or if it’s a pleasure to read.
Chasing a perfect MTBF is a noble but ultimately hollow pursuit if the experience it safeguards is allowed to crumble. Reliability is a holistic quality, a symphony of interdependent parts, not a single, silent, unmoving hand on a dial. The most reliable service isn’t the one that never goes down; it’s the one whose internal state is so well understood, whose subtle changes are so clearly heard, that we can ensure its true purpose—serving users effectively—never truly fails.
Notes & further reading
A few pages I came back to while writing this:
- Madison, WI
- The Sentinel's Sway: On the Fallacy of Perfect Posture
- Milwaukee, WI
- The Unwound Metronome: On the Wisdom of Irregular Heartbeats
- a useful directory
- The Gardener's Silent Sprinkler: On the Reliability of an Unseen Nourishment
- a local resource
- a place-by-place guide
- one area's overview
- a regional guide
- a helpful reference
- a practical rundown
- a nearby resource