The Signalman's Broken Block: On the Delusion of the Intermittent Path

Before we spoke of digital pings and API endpoints, reliability was enforced by mechanical certainty. A stark example lies in the tracks of the early railways, governed by a system of interlocking signals and tokens designed for an uncompromising rule: only one train could occupy a single, defined section of track, known as a "block." The signalman’s duty was absolute, his instruments a physical manifestation of state. A lever thrown in one signal box would lock the corresponding levers in adjacent boxes. A clear signal meant the block was, without question, clear.

But what of the intermittent fault? The telegraph wire that passed the "all clear" message would occasionally fray, its copper core exposed to the elements. In fair weather, the circuit would close, and the signal would pass through. A single drop of rain, however, could bridge the gap between the broken wire and a nearby ground, shorting the circuit. To the signalman at the receiving end, the telegraph would click to life, indicating the block ahead was empty. But the signal was a lie, born of a transient condition he could not see. The message was not "clear"; it was the ghost of a connection, a deceptive spark mimicking the real thing.

This historical vignette is a chillingly precise analogue for a modern horror: the flapping network route or the service that passes its health checks 99 times out of 100. Our automated systems, much like that 19th-century signalman, rely on a steady stream of truths. A health check is a question posed to a service: "Are you well?" A single, passing response can be that shorted telegraph wire—a momentary alignment of cosmic luck that masks a deeper, persistent instability. The service might be choking on memory leaks, its database connections might be fraying, but for the exact half-second our probe arrives, it gasps a convincing "yes."

The danger is not the constant failure. A service that is consistently down declares its own inadequacy. The true peril lies in the illusion of health, the intermittent path that convinces us the track is clear when a catastrophe is merely waiting for the right conditions—the equivalent of a fully-loaded express train entering the block. We are lulled by a cadence of successful pings, mistaking a statistical probability for a guarantee of integrity. We become the signalman who, after a hundred false clears, trusts the broken instrument more than the ominous silence that preceded it.

The lesson from the rails is that reliability cannot be built on a single, simplistic signal. It requires a multi-faceted understanding of state. The signalmen, in their wisdom, developed a system of physical tokens and mechanical interlocks that could not be fooled by a spurious telegraph signal. Similarly, our observability must extend beyond a binary "up/down." We need the equivalent of train journals, track-side observers, and pressure gauges on the steam lines. We need logs, traces, and metrics that reveal the internal state, the resource pressure, the gradual degradation that a simple health check will never see. The goal is not just to know if the service is momentarily responsive, but to understand the true condition of the track ahead, ensuring that the path is not just intermittently open, but fundamentally sound.

Notes & further reading

A few pages I came back to while writing this: