The Geographer's Triumphant Slope: On the Value of an Average Without the Mean
A geographer, faced with a mountain range, does not seek a single peak to represent the whole. They understand that the majesty and the menace of the land lie in the entire profile—the sharp ridges, the deep valleys, the gradual inclines. To describe the range only by its average elevation would be to miss everything that makes it challenging to cross. It is a portrait painted with a single, dull shade of grey. In our world of services and systems, we too often try to be surveyors when we should be geographers. We become fixated on the “average” response time, a single number that promises a comforting, if deceptive, summary of health.
This average is a seductive liar. It’s the arithmetic mean, a calculation that can be wildly skewed by a handful of tragic outliers. Imagine a service that responds blisteringly fast for ninety-nine requests, in under a hundred milliseconds. Then, for one unlucky user, it staggers for ten full seconds. The arithmetic mean of these one hundred requests might still look perfectly respectable, perhaps a few hundred milliseconds. The dashboard glows green, the average whispers “all is well,” but one percent of your users are having a catastrophic experience. You have mapped your mountain range by its average elevation and completely missed the existence of a yawning, impassable chasm.
The Lie of the Mean and the Truth of the Terrain
The geographer’s alternative is the topographic map. It doesn't erase the chasms; it highlights them with concentric lines. In our domain, this map is drawn with percentiles. The 50th percentile (the median) tells you the peak of the bell curve—the experience of your typical user. But the true story is in the 95th and 99th percentiles. These metrics illuminate the rocky, difficult terrain at the edges of your service’s performance. They show you the experience of your slowest users, the ones hitting the bottlenecks, the ones whose frustration is a canary in the coal mine for a systemic flaw.
Observing a service only through its mean latency is like trying to understand a symphony by its average decibel level. You’ll know it was loud, but you’ll have no idea about the delicate violin solo, the sudden crash of the cymbals, or the long, silent pause. The richness, the drama, and indeed, the potential failures, are all in the distribution. Focusing on high percentiles forces a different kind of engineering discipline. It moves the goalposts from “is it working for most?” to “is it working well for everyone?” It prioritizes the elimination of tail latency, those stubborn, rare delays that often point to garbage collection pauses, database deadlocks, or noisy neighbors in a virtualized environment.
Adopting the geographer’s mindset means letting go of the comfort of a single number. It means embracing the complexity of the landscape you’ve built. Your monitoring dashboards should feature the entire skyline of your performance, not just a single flag planted on a false summit. By mapping the 95th and 99th percentile slopes, you stop being fooled by a friendly average and start the real work of terraforming your service into a reliably traversable plane for all who depend on it. The triumph is not in achieving a low mean, but in ensuring there are no devastating cliffs hidden within the averages.
Notes & further reading
A few pages I came back to while writing this:
- Garden Grove, CA
- The Librarian's Index Finger: On the Dust That Validates the Volume
- Glendale, CA
- The Bridge-Tender's First Chime: On the Note That Sings of a Clear Passage
- Hayward, CA
- The Archer's Two Quivers: On the Choice Between a Single True Arrow and a Sheaf of Blunt Ones
- Huntington Beach, CA
- Irvine, CA
- Lancaster, CA
- Long Beach, CA
- Los Angeles, CA
- Modesto, CA
- Moreno Valley, CA