The Sculptor's Lost Thumb: On the Tyranny of the Perfect Availability Score

In the quiet, obsessive world of the sculptor, the aim is perfection. The final work should be seamless, a flawless representation of the artist's intent. We in the world of running online services have adopted a similar, almost artistic obsession: the pursuit of a perfect availability score. Three, four, or even five nines of uptime have become our marble statues, polished to a gleaming finish and displayed as the ultimate testament to our craft. We tend to our dashboards, watching the glowing green lights and the unwavering 99.999% with a sense of deep satisfaction. But what if this gleaming statue is a lie? What if, in our quest for perfect uptime, we are losing something more vital, like the sculptor who, in smoothing away every tiny flaw, removes the thumb that gives the hand its character and strength?

This is the tyranny of the perfect score. It is a received wisdom so deeply ingrained that to question it feels like heresy. But it’s a metric that, in its singular focus, can create a culture of fear and fragility. When the only thing that matters is that the green light stays green, the natural, healthy process of a service experiencing minor, recoverable failures is driven underground. Teams become terrified of deploying changes, even beneficial ones, because a momentary blip will tarnish the pristine record. Innovation slows to a crawl, sacrificed on the altar of an unblemished percentage point. The system becomes a museum piece—perfectly preserved, but no longer evolving.

Worse still, this obsession often blinds us to the user’s actual experience. A service can technically be "up" while being functionally unusable. Latency can spike to agonizing levels, partial failures can corrupt data silently, or a critical feature can fail while the main health-check endpoint continues to return a cheerful "200 OK." The perfect availability score becomes a hollow victory, a Potemkin village hiding a landscape of user frustration. We celebrate not having a catastrophic outage while ignoring the thousand small paper cuts that drive people away daily. We are monitoring the sculpture from a distance, admiring its silhouette, but failing to notice the hairline cracks spreading across its surface.

True resilience isn't the absence of failure; it's the capacity to fail gracefully and recover quickly. A system that has never experienced a failure in production is an untested system, a ticking clock. The wiser approach is to build an observability practice that understands the spectrum of health. It means measuring not just uptime, but also latency distributions, error budgets, and the speed of our recovery. It means embracing controlled chaos, like deliberate game days where we learn to navigate failures in a safe environment. It’s about building systems that are like a seasoned forest: they experience storms, lose branches, but their deep roots allow them to recover and regrow stronger.

Perhaps it's time to put down the polishing cloth and pick up a different tool. Instead of a tyrannical pursuit of a perfect, sterile uptime score, we should aim for a system that is robust, observable, and humane. A system that allows for the occasional stumble because it has the strength to get back up, learning from the experience. A perfect statue is impressive, but a living, breathing, adaptable system is what truly serves its purpose. The thumb might not be perfectly sculpted, but it’s what allows the hand to grasp, to create, and to feel. And that is far more valuable.

Notes & further reading

A few pages I came back to while writing this: