The Lighthouse Keeper's First Mistake: On the Dangers of an Unwritten Log
There’s a story, perhaps apocryphal, whispered among maritime historians about a lighthouse keeper on a notoriously rugged stretch of the Atlantic coast. He was diligent, dedicated to his flame. Every night, without fail, he would light the great lamp, polish the lenses, and watch the beam cut through the squalls. By all observable metrics, his service was supremely reliable. Ships passed safely, their captains comforted by the unwavering light.
But one autumn, a ferocious storm rolled in. The keeper performed his duties perfectly. The light burned bright. Yet, a schooner, trusting in that familiar beacon, foundered on a hidden reef. The tragedy was a mystery. The light had never faltered. It was only later, in the investigation, that the keeper’s critical oversight was discovered. In his logbook, the entries for the preceding weeks were sparse. He had noted the lamp’s operation, but he’d omitted a smaller, seemingly trivial detail: the grinding noise from the clockwork mechanism that rotated the lens had changed its pitch. It had become slightly slower, a deeper groan.
The Unobserved Failure Mode
This subtle change was the canary in the coal mine. A gear was wearing down, imperceptibly slowing the rotation of the beam. To a ship’s captain far out at sea, this didn’t manifest as a flicker or a failure. The light was still there, bright as ever. But its timing was off. The interval between flashes, the unique signature that identified this specific lighthouse, had elongated. The ship’s navigators, working from old charts, were expecting the light at a specific rhythm. The slowed beam created a deadly illusion, misjudging the ship’s position by a few critical hundred yards—just enough to meet the reef.
In our world of digital services, we have our own great lamps: the HTTP status codes, the ping responses, the green dashboards that tell us a service is "up." We religiously monitor these beams. But like the keeper, we can fall into the trap of monitoring only the most obvious signal. Our service might return a 200 OK, but what about the 95th percentile latency creeping up by 50 milliseconds? What about the slowly increasing error rate on a dependent API that hasn’t yet tripped a failure threshold? These are the changes in pitch, the grinding noises of our systems.
The keeper’s first mistake wasn’t the failing gear; it was the failure to record the anomaly. He heard the sound, dismissed it as insignificant, and never wrote it down. He lacked the practice of observability—the discipline of capturing not just the binary state of "working" or "broken," but the rich tapestry of internal state and nuanced signals that tell the full story of a system’s health. Without those logs, that telemetry, he had no way to correlate the eventual disaster with its early, quiet precursor.
The lesson is stark. Uptime is not the same as correctness. A service can be up and yet be profoundly wrong, leading its consumers quietly astray. True reliability demands that we listen for the grinding gears, not just watch the light. It requires us to log the subtleties—the timings, the partial errors, the slight deviations—because today’s anomalous log entry is tomorrow’s root cause analysis. The keeper’s logbook, had it been complete, would not have prevented the gear from wearing, but it would have provided the context needed to understand the threat before it was too late. It’s a reminder that in our pursuit of reliability, the most dangerous entry is the one we never make.
Notes & further reading
A few pages I came back to while writing this:
- one area's overview
- The Barometer's Sticky Needle: On the Calm Before the 404
- a practical rundown
- The Unlit Wick: On the Folly of an Unobserved Candle
- Little Rock, AR
- The Potter's Centering Hand: On the Stability Before the Spin
- Gilbert, AZ
- Peoria, AZ
- Surprise, AZ
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT
- Stamford, CT