The Tuning Fork's Hum: On the Perfect Pitch for a Health Check Interval

We set up our health checks, those faithful sentinels, and we give them a number: check every 30 seconds. Or 60. Or 5. It’s often an inherited setting, a piece of tribal knowledge, or a round number that feels right. But that number—the interval—is the fundamental frequency at which your service hums to itself, a constant question asked into the void. Getting it wrong doesn’t just waste resources; it creates a kind of observational noise that either masks failure or manufactures it. The trick isn’t to find the ‘best’ interval, but to tune yours to the specific resonance of the service it guards.

The concrete technique is simple, yet I see it overlooked: graph your service’s natural failure and recovery duration, then set your check interval to be shorter than the former and longer than the latter. You need two histograms. First, track how long your service is typically in a genuinely failed state before human or automated intervention fixes it. This is your ‘Mean Time To Repair’ (MTTR) in the wild. Second, and this is critical, track how long your service’s own startup or self-healing routines take. That’s the time from when the process starts to when it’s truly ready to serve.

The Space Between Notes

Once you have these two numbers, the logic unfolds. Your check interval must be shorter than your typical failure duration, or you risk missing failures entirely. If a database hiccup resolves itself in 90 seconds on average, a 120-second check might blissfully skip over the whole event. But more subtly, your interval must be longer than your service’s recovery time. If your application takes 45 seconds to warm up its caches and accept connections, a 30-second check will perpetually see a ‘starting’ service as ‘down,’ triggering false alarms with every deployment or restart. You’ve tuned your monitor to a note that your service cannot hold.

This tuning creates a necessary and healthy buffer—a silence between the questions that allows the system to breathe, to fail meaningfully, and to recover in peace. For a stateless API that starts in 2 seconds, a 10-second check is aggressive but sane. For a legacy monolith that grumbles to life over a 3-minute span, you must grant it that grace. Your check interval is not a measure of your vigilance, but of your understanding. It says, “I know how you break, and I know how you heal, and I will listen at a rhythm that respects both.”

So, open your metrics. Plot the recovery times from your last ten deploys or restarts. Look at the duration of last month’s transient failures. The right number will emerge from this data, not from a convention. It will be an odd number, perhaps 23 seconds or 117 seconds. It will be your service’s unique pitch. When you strike it, the hum will not be one of anxious polling, but of harmonious alignment—a signal clean enough to trust in the silent moments between the checks, when the real work is done.

Notes & further reading

A few pages I came back to while writing this: