The Clockmaker's Regret: On the Tyranny of the Average Response Time
There's a story, perhaps apocryphal, about a clockmaker who prided himself on precision. To prove the accuracy of his latest masterpiece, he placed two clocks in his shop window, synchronized to the second. A week later, a patron pointed out a problem: one clock was a minute fast, the other a minute slow. The clockmaker, unfazed, declared them perfectly accurate. "On average," he explained, "they tell the exact time."
Our reliance on the average response time metric in service monitoring is a bit like that clockmaker's flawed logic. It’s a comforting number, a single, digestible figure that we can track on a dashboard, chart over time, and report in status meetings. We strive to lower it, to keep it in the green. But this singular focus on the mean, the arithmetic middle ground of all our request durations, can become a tyrant. It creates a placid, average reality that obscures the jagged, often distressing truth of the user experience.
Consider what the average hides. A service might have a very healthy-looking average response time of 200 milliseconds. This number could be composed of 99 requests that each took a zippy 150ms, and one single, unfortunate request that languished for a full five seconds. Mathematically, the average is indeed around 200ms. But for that one user whose transaction took five seconds, the service was essentially broken. Their experience wasn't "average"; it was terrible. They are the customer who abandons their cart, refreshes the page in frustration, or writes a scathing review. The average, in its desire for a tidy summary, has completely erased their pain.
The Ghosts in the Machine
This is where observability must step in to challenge the tyranny of the average. The outliers, the long tails of the latency distribution, are not statistical noise to be smoothed away. They are the ghosts in the machine, the whispers of a deeper problem—a slow database query under load, a caching layer miss, a garbage collection pause, a resource contention issue in a neighboring service. By focusing solely on the average, we are telling our monitoring systems to ignore these critical signals. We are asking them to prioritize the comfort of the majority over the crucial diagnostic data represented by the few.
The remedy isn't to discard the average, but to dethrone it. It should be one voice in a choir, not the sole soloist. We must learn to listen to the percentiles—the 95th (p95), the 99th (p99), and even the 99.9th (p999). These metrics tell us about the experience of our slowest users. They answer the question, "How bad is it for the users having the worst time?" A growing gap between your average and your p95 latency is a blazing alarm bell, signaling that a subset of your traffic is encountering a problem that the majority is blissfully unaware of.
Monitoring a service for reliability is not about ensuring that most requests are handled well. It is about ensuring that virtually all requests are handled acceptably. It requires a vigilance that seeks out the exceptions, the anomalies, the struggles at the edges. We must move beyond the clockmaker's regret of seeing only the harmonious average, and instead develop the keen ear of a mechanic, listening for the faintest knock or whir that betrays a future breakdown. For in the realm of reliable services, it is the experience of the outlier that truly defines your system's health.
Notes & further reading
A few pages I came back to while writing this: