The Potter's Thrown Gauge: On Measuring the Spin, Not Just the Crack
We spend a great deal of time listening for the crack. A server fails to respond, an API endpoint returns a 500 error, a database connection times out. These are the sharp, audible cracks in our systems, the unmistakable signals of a failure. Our monitors scream, our alert channels light up, and we scramble. This, we tell ourselves, is observability—the art of detecting when the vase has broken. But what if we could understand the moment the flaw was introduced, long before the firing, when the clay was still spinning on the wheel?
There’s an old tool in the potter’s workshop, less celebrated than the kiln or the glaze, called a thrown gauge. It doesn’t measure the final product. Instead, it’s a simple arm with a needle that the potter gently rests against the spinning clay. The goal isn't to see if the clay is cracked; it’s to measure the wobble, the subtle, almost imperceptible unevenness in the rotation. A perfectly centred piece spins true, its surface a smooth, humming blur. An off-centre piece shudders, its imbalance telegraphing a future of structural weakness. The potter adjusts their hands immediately, correcting the spin. They are not fixing a break; they are preventing the conditions that lead to one.
Our digital services are no different. Our standard uptime checks and health endpoints are the kiln’s temperature alarm—they tell us when disaster has already struck. They are vital, but they are fundamentally reactive. The real art of running reliable services lies in measuring the ‘spin’ of our systems while they are still running. This is the domain of request latency, but not as a single, monolithic number to be checked against a threshold.
A true thrown gauge for a web service looks at the distribution of latency across percentiles. The average response time, the P50, might be a comfortable 200 milliseconds, spinning smoothly. But what about the P95, or the P99? These are the wobbles. A creeping increase at the 99th percentile, from 800ms to 1200ms, isn’t a crack. The service is still ‘up’ by every conventional measure. Yet, that wobble is a critical signal. It’s the first sign of resource contention, a memory leak beginning to pool, a downstream dependency starting to strain. It’s the moment the clay drifts a millimetre off-centre.
By the time the P99 latency degrades into a full timeout—a crack—the failure mode is set. The incident is live. We are no longer the potter at the wheel; we are the restorer in the museum, piecing together fragments with glue. Observability, in its deepest sense, is about installing thrown gauges throughout your architecture. It’s about having the granular, real-time telemetry to see the wobble in your database queries, the shudder in your message queues, and the uneven spin of your caching layer. It’s a shift in mindset from simply ensuring the wheel is turning to constantly feeling for the true, balanced centre of its motion. The goal is not just a service that doesn't break, but one that spins so flawlessly you can almost forget it’s there at all.
Notes & further reading
A few pages I came back to while writing this:
- one area's overview
- The Archivist's Dust-Free Ledger: On the Integrity of the Unchanged Record
- a practical rundown
- The Weaver's Loom and the Cobweb: On the Tension Between Structure and Emergence
- Little Rock, AR
- The Electric Kettle's Quiet Click: On the Final State of Readiness
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT
- Stamford, CT