The Night Watchman's First Round: On the Humble Power of the Baseline Check
Every service has a sound. It’s not a literal hum or buzz, but a distinct digital cadence composed of its response times, its error rates, the gentle rhythm of its resource consumption. When it’s healthy, this sound is a quiet, consistent melody in the background of your infrastructure. When it begins to falter, the melody develops a dissonance long before it becomes a clamor. The challenge, for those of us tasked with keeping the lights on, is learning to hear that dissonance the moment it begins.
We often rush to sophisticated observability platforms, dashboards bristling with a hundred metrics, convinced that complexity is the key to understanding. But I’ve found that the single most powerful technique for knowing when something is truly ‘off’ is also the simplest: establishing a performance baseline. This isn’t just a number; it’s the foundational truth of your service’s normal behavior, its unique signature when left to its own devices under typical load.
The how-to is straightforward, and you can start with just one critical service. First, choose your primary metric. For most web services, the 95th percentile latency is a superb candidate, as it ignores outliers and focuses on the experience of the vast majority of your users. Next, you need a period of known stability. This is the crucial part. Pick a 24-hour period—a regular Tuesday is perfect—when you are confident the service was running smoothly, with no major deployments, no traffic spikes from a marketing campaign, and no underlying infrastructure issues. This day becomes your ‘normal.’
Now, calculate the average and standard deviation for your chosen metric over this period. Don’t just take a single average for the whole day; you’re smarter than that. A service behaves differently at 3 AM than it does at 3 PM. Break it down. Calculate a baseline for each hour of the day. Your 2 PM baseline might be 180ms ± 20ms, while your 3 AM baseline is 110ms ± 5ms. This gives you a dynamic, time-aware understanding of health.
Here’s where the technique transforms from a static number into a living guardrail. Configure your alerting not on a single, arbitrary threshold (like “alert if latency > 500ms”), but on deviation from this learned baseline. Set an alert to trigger if, for example, the current latency exceeds the baseline for that hour by more than two standard deviations. This simple change is transformative. It means you’re no longer alerting on ‘bad’ in an abstract sense, but on ‘abnormal.’ It accounts for the natural ebb and flow of your system’s life.
This baseline becomes your night watchman’s familiar route. The watchman doesn’t just check if doors are locked; he knows the precise sound of the floorboards, the exact shadow a certain lamppost casts. A change, however slight, is immediately apparent. By defining normal with this granularity, the signal of a genuine problem emerges from the noise of daily operation with startling clarity. It’s the whispered question answered before the alarm bell needs to ring, the detection of the subtle warp long before the pot cracks.
Notes & further reading
A few pages I came back to while writing this: