The Apothecary's Rubbing Aloes: On the Unpleasant Truth of a Synthetic Distress

Every reliable service is a carefully managed garden, full of green indicators and the steady hum of healthy processes. We tend to it with our dashboards, our heartbeats, and our pings. But there’s a quiet, uncomfortable truth in this craft: we often monitor for comfort, not for reality. We check to be reassured, not to be challenged. What we need, on occasion, is not another gentle probe but a deliberate, synthetic distress—a controlled dose of unpleasantness to verify our remedies actually work.

Think of it as the apothecary’s method. For centuries, an apothecary’s stock of dried aloe was considered inert, a mere ingredient on a shelf. Its true potency wasn’t known until it was rubbed, a deliberate friction that released its curative sap. Our alerting systems, runbooks, and escalation chains are that dried aloe. We document them, we review them, but their true state—their ability to actually relieve a problem—remains theoretical until they are put under the specific friction of a real, but safe, failure.

The technique is simple, yet profound: regularly schedule a synthetic incident that triggers your real alerting pipeline. This is not a chaos engineering experiment that takes down a database. It’s far more surgical and mundane. You engineer a single, specific check to fail in a way that mimics a genuine, mid-tier problem. Perhaps your health check endpoint begins returning a 503 after a one-minute delay. Maybe a synthetic transaction designed to test the checkout flow begins to log a cryptic, but unique, error code at a 10% rate. The key is that the symptom must be plausible, non-catastrophic, and must travel the exact same path as a real alert.

Then, you watch. Does the page go to the right team, or is the routing stale? Does the alert include the necessary context, or is it a barren, panicked string of numbers? Does the on-call engineer’s runbook link still resolve? When they follow its first diagnostic step—say, to check a specific metric or log pattern—is that tool actually accessible with their current permissions? The synthetic distress isn’t testing your systems; it’s testing the connective tissue between your systems and your people. It’s rubbing the aloe to see if any sap appears.

The value of this friction is not in finding that your paging system works. It’s in discovering the thousand paper cuts that have accumulated since your last real outage: the deprecated dashboard, the new team member never added to the rotation, the runbook step that references a server decommissioned six months ago. It turns your response plan from a static document into a living, slightly uncomfortable, but far more truthful practice. You stop monitoring for a clear sky, and start practicing for the specific texture of the coming storm. You learn not just if your line is humming, but if you can still hear it clearly when it begins to crackle.

Notes & further reading

A few pages I came back to while writing this: