The Astrologer's Perfect Ephemeris: On the Deceptive Order of P95 Latency

In the ancient pursuit of order among the stars, an astrologer’s ephemeris tables predict the celestial dance with mathematical precision. Each planet, a predictable point of light, follows a prescribed path. For those of us building modern services, our version of this table is often the latency chart, and its most revered planetary body is the 95th percentile, or P95. We plot its course, we track its movements, and we proclaim our system’s health by its steady, predictable orbit. But what if our faith in this single number is a form of astronomical hubris? What if, in our quest for a tidy summary, we are ignoring the cosmic chaos that truly defines our service’s universe?

P95 is not a lie, but it is a powerful simplification. It tells us, reassuringly, that 95 out of 100 requests met a certain speed threshold. It gives us a benchmark for performance that feels concrete, a number we can trend, alert on, and report upwards. It allows us to say, "Our P95 is under 200ms," with the confident air of an astrologer predicting a planetary transit. Yet, this comfort is precisely the danger. The P95 metric draws a neat circle around the vast majority of our requests, leaving the crucial 5%—the outliers—dismissed as statistical noise, cosmic anomalies unworthy of serious charting.

This is the astronomer’s folly: focusing so intently on the predictable planets that one misses the comets, the supernovae, and the gravitational anomalies that rewrite the rules of the cosmos. Those last five requests out of a hundred are not noise; they are the story. They are the user in a remote geographic region, struggling against a poor network path our monitoring doesn’t probe. They are the cache miss that cascades into a thundering herd problem under a specific, un-modeled load pattern. They are the database query that, just once in a blue moon, decides to perform a full table scan. To treat these as mere rounding errors is to ignore the very failures that erode user trust most severely.

A P95-centric worldview creates a system that is optimized for the middle of the distribution, potentially at the expense of the tails. An engineering team, pressured to keep that P95 line flat, might implement a change that helps 95% of requests shave off 10 milliseconds, while inadvertently doubling the latency for the remaining 5%. The dashboard glows green, the reports are filed, but a small, yet significant, cohort of users now experiences a service that feels broken. Their reality is invalidated by our aggregate success.

So what is the alternative? It is to become less of an astrologer and more of an astronomer—or better yet, an astrophysicist. It is to treat P95 not as the definitive answer, but as the starting point for a deeper inquiry. We must turn our telescopes to P99, P99.9, and even higher percentiles. We must embrace the chaos of the tail by looking at the full, un-summarized distribution of requests through tools like heatmaps or histograms. We must correlate these latency outliers with other signals—error rates, infrastructure metrics, business events—to understand the 'why' behind the anomaly. The goal is not to eliminate the P95, but to recognize it as one star in a much larger and more turbulent constellation. The truest measure of a system’s reliability is not how well the majority of requests behave, but how gracefully we handle the exceptions.

Notes & further reading

A few pages I came back to while writing this: