The Navigator's Fixed Star: On the Tyranny of a Constant Bearing

We are taught to revere consistency. In the realm of service reliability, this reverence manifests as a near-obsession with the flat line. A perfect, unwavering 99.999% uptime graph is our Sistine Chapel ceiling, a thing of beauty and a goal to be worshipped. We set our alerts to scream at the slightest deviation, chasing the phantom of a perfect, unbroken service. But in our quest for this constancy, I fear we have mistaken the map for the territory, the bearing for the journey itself.

The received wisdom is simple: a straight line is good, a dip is bad. Our monitoring dashboards are designed to reinforce this, with green lines stretching infinitely to the right, a testament to our engineering prowess. We celebrate these periods of flawless operation, and rightly so. But this focus on the constant bearing can blind us to a more subtle, and perhaps more important, truth. A service that is merely ‘up’ is not necessarily ‘well’. It can be limping, degraded, or operating in a state of immense internal strain, all while our primary health check continues to return a triumphant 200 OK.

This is the tyranny of the fixed star. The ancient navigators knew that steering solely by a single, unmoving point could lead you onto the rocks if you failed to account for the currents beneath you and the drift of your own vessel. In the same way, our singular focus on binary uptime can mask the underlying currents of growing latency, increasing error rates within successful responses, or diminishing throughput. The service is technically afloat, but it is no longer sailing well; it is being dragged.

True observability—the real art of navigation—requires us to look beyond the horizon line of ‘up’ or ‘down’. It demands we listen for the creak of the rigging in the form of application logs, feel the pull of the current through distributed tracing, and constantly take soundings with synthetic transactions that test entire user journeys, not just a single endpoint. A health check that only verifies a heartbeat is like a navigator who only checks that the ship hasn’t sunk. It’s a low bar.

The goal should not be a flat line, but a predictable and understood rhythm. A healthy service has a pulse, not a steady tone. It has ebbs and flows of traffic, minor, expected variations in performance, and a known capacity for strain. The frightening graph isn’t the one with a small, handled dip; it’s the one that is deceptively flat right up until the moment it falls off a cliff. By worshipping the constant, we stop learning the language of our systems. We trade the nuanced story of a living service for the bland headline of a persistent one. Let us navigate by the whole sky, not by a single, misleading star.

Notes & further reading

A few pages I came back to while writing this: