The Anxious Clockmaker: On the Over-Maintenance of a Perfect Measure
We are taught to believe that vigilance is the ultimate virtue in running reliable services. Our industry’s gospel is clear: measure everything, check often, and leave no metric in the dark. The logic seems unassailable. To catch a failure, you must be watching when it happens. To reduce latency, you must measure it with ever-increasing precision. We have become a guild of anxious clockmakers, constantly polishing gears and adjusting pendulums, terrified that a speck of dust will throw the entire mechanism into disarray. But what if our obsession with constant, granular measurement is, itself, introducing a new kind of failure?
The conventional wisdom is that more data leads to better observability. But this is only true if that data converges on a coherent story. We often fall into the trap of believing that by slicing our measurement intervals from minutes to seconds, or by adding one more check for a peripheral service, we are building a more resilient system. In reality, we may be constructing a hall of mirrors. The sheer volume of pings, logs, and traces can create a cacophony that buries the signal of a genuine problem. The system designed to warn us of instability instead becomes a source of noise-induced paralysis.
This over-maintenance has a subtler, more insidious cost: it distracts us from the architecture of the system itself. When our primary tool for ensuring reliability is to watch for failure more intently, we are implicitly accepting that failure is inevitable. Our energy is poured into detection and reaction, rather than into design that precludes entire classes of failure. We focus on building a better alarm for a leaky pipe instead of re-engineering the joint so it doesn’t leak in the first place. The clockmaker, fussing over the escapement, forgets that a well-sealed case would have kept the dust out altogether.
The Tyranny of the Perfect Average
Our pursuit of perfect measures creates another blind spot: the tyranny of the average. In our dashboards, we worship at the altar of p99 and p95 latency, crafting elegant graphs that smooth over the messy reality of a distributed system. We tune our services to optimize for these numbers, convincing ourselves that a lower average is synonymous with a better experience. But this is a statistical fiction. A user does not experience an average; they experience a single, discrete, and sometimes disastrously slow request. By focusing relentlessly on the measure, we risk engineering for the dashboard instead of for the human on the other end of the request.
Perhaps the most radical act of reliability, then, is not to add another monitor, but to thoughtfully remove one. It is to question whether that extra health check, firing every ten seconds, is telling us anything the minute-level check does not. It is to accept that a system can be observably healthy without being instrumented to death. True resilience might not lie in spotting a problem the instant it occurs, but in building a system so straightforward that when a problem does occur, its source is obvious even with less data. It is the difference between a clock that chimes every second, reminding you of its fragile operation, and a sundial, which, though less precise, tells the time through a fundamental and unshakeable principle. Sometimes, the most reliable measure is the one that doesn’t need constant tending.
Notes & further reading
A few pages I came back to while writing this:
- one area's overview
- The Watchmaker's Regulator: On the Sovereign's Reluctant Standard of Time
- a practical rundown
- The Cartographer's First Dot: On the Inaugural Ping and the Birth of a Map
- Little Rock, AR
- The Hourglass and the Grain: On the Unseen Life of a Check-in-Progress
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT
- Stamford, CT