Alert Actionability Lessons From an Emergency Room Study
An emergency department study found 0.8 percent of monitor alarms changed clinical management. How to measure alert actionability on your own on-call.

Quick Answer
An emergency department study found that only 0.8% of monitor alarms changed clinical management, and staff did not overtly respond to 64% of them. Software teams that celebrate fewer pages are measuring volume, not alert actionability. Measure how often a page changes what the responder does, track time to first action, and expect the pages that survive filtering to be harder.
Engineering leaders often celebrate when an alert hygiene project cuts page volume by half. It is a clean number to show a VP of Engineering. We think it hides a harder question, which is if the pages that remain change what the responder does.
Alert actionability is the share of alerts that lead to a different action than the responder would otherwise have taken. Emergency medicine measured it directly, and the result is a useful mirror for on-call engineering.
The Number You Can Show Leadership and the Number That Matters
When you filter for actionable alerts, you remove the trivial ones. What remains are the cases where a correlation engine hands a tired engineer three plausible causes at once. We think filtering does not remove judgement, it concentrates it.
The win leadership sees, fewer interruptions, can hide a new debt. The cognitive load of testing hypotheses in the middle of the night rises as the easy pages disappear, and no dashboard shows it.
Emergency Medicine Ran the Study
Software operations is not the first field to struggle with alert fatigue. Fleischman and colleagues published an observational study in the American Journal of Emergency Medicine in 2020. A physician observer watched an urban academic emergency department for 53 hours and recorded each monitor alarm, how staff responded and if it changed clinical management.
The abstract reports 1,049 alarms for 146 monitored patients, a median of 18 alarms per hour. Only 8 of the 1,049 led to changes in clinical management, which is 0.8% with a 95% confidence interval of 0.3% to 1.3%, and those changes occurred in 5 of the 146 patients. The AHRQ summary notes that staff did not observably respond to nearly two-thirds of alarms, which may be a sign of alarm fatigue.
We value the design because it does not rely on self-reported feelings of burnout. It records a binary outcome by direct observation, which is if the signal changed what the human did.
Why 0.8 Percent Is Not the Shocking Part
The more striking finding is the staff response. Staff did not respond overtly to 64% of alarms, according to the abstract. We read that as a rational adaptation to a system that mostly cries wolf, and not as individual carelessness.
Medicine had to stop measuring noise and start measuring change in action, because the stakes are high. Software on-call often shares the same always-on criticality and has not yet adopted the same rigour.
Key Takeaway. The only metric that justifies an alert is if it changed the responder's subsequent action.
Measuring Alert Actionability With a Two-Week On-Call Audit
To move past volume reporting, we would run an actionability audit over a two-week window. It should be a blameless quality project and not a performance review.
- Capture in real time. For every page, the responder records four things immediately, because recall afterwards is unreliable. These are the trigger, the first dashboard or command checked, the first corrective action and if the alert changed what they were going to do.
- Compute the action rate. Divide the number of action-changing alerts by the total number of alerts.
- Measure time to first action. Track the delta between the page firing and the first diagnostic command.
If responders think this is a test of their efficiency, they will inflate the numbers. Be explicit that a low actionability rate is a failure of the monitoring system and not of the engineer. Google's SRE book states the principle directly in its chapter on monitoring distributed systems. "Every page should be actionable," it says, and pages with rote, algorithmic responses should be a red flag.
What Medicine Did Next With Alarm Thresholds
The study did not simply recommend fewer alerts. Its authors suggested expanding alarm thresholds, customising parameters to the patient's clinical status and monitoring more selectively.
In software, we translate that into moving away from global thresholds such as alerting when the 5xx rate exceeds 1% and toward state-aware monitoring. An alert should depend on the service's context, such as its recent deployment history and current traffic. Deduplication and smarter routing help, but we see them as presentation layers that organise the noise without making the signal more actionable. Our guide to reducing on-call alert fatigue by triaging first covers the tuning in practice.
The Second-Order Cost Nobody Budgets
When you eliminate non-actionable alerts, you raise the average difficulty of the rotation. The remaining alerts are rarer and less familiar, so engineers cannot build muscle memory from frequent simple fixes.
Grouping signals is a presentation and not a diagnosis. A tool that says ten signals fired together helps, and it still leaves a hypothesis elimination task. If nobody budgets for it, senior engineers burn out on low-volume rotations because the three pages they got needed four hours of forensics each. We would give every surviving alert three things.
- An attached runbook with the first three diagnostic commands.
- A dependency model that separates upstream causes from downstream effects.
- Explicit recognition that a grouped alert is a request for investigation and not a simple restart.
The investigation after filtering is cross-layer work across code, infrastructure and logs. It is the part where we would use dedicated AI agents that investigate with cited evidence, while the responder decides.
What to Report Upward About Alerting
To give leadership a true picture, we would report three numbers together.
- Pages per rotation. The usual volume metric, ideally trending down.
- Actionability rate. The percentage of pages that changed a responder's behaviour, ideally trending up.
- Time to first diagnostic action. How quickly an engineer can orient, ideally flat or falling.
If volume falls and time to first action rises, you have concentrated judgement. You filtered the noise, and you now owe the team diagnostic support for the complexity that remains. Our notes on escalation metrics and answerability rate apply the same discipline to escalations, and on-call onboarding and time to first page covers the human side of the rotation.
What an Actionability Audit Cannot Tell You
An emergency department is not an on-call rotation, and we would not stretch the analogy. Monitor alarms and pages differ in cost and context, and a software page that confirms a responder's suspicion may still be worth its place.
The audit also measures the past. A page that never changes behaviour might be guarding a rare catastrophic failure, and removing it on the numbers alone would be a mistake. We would treat a low action rate as a prompt for review of that alert, and not as an automatic deletion.
Measure the Change in Action, Not the Count
Alert actionability is the metric that shows if on-call is working. Run the two-week audit, report the three numbers together and budget diagnostic support for the harder pages that remain.
We would start with the five noisiest alert rules. If you want agents to help with the hypothesis work once the noise is filtered, Operate runs dedicated AI agents that investigate production problems with cited evidence and propose fixes for your team to review.
Frequently Asked Questions
We define it as the share of alerts that lead a responder to a different action than they would have taken otherwise. It measures change in behaviour, not volume, which is why a project can cut pages in half and leave actionability unchanged.
We see it when most alerts do not need action, so responders adapt by ignoring many of them. The emergency department study found staff did not respond overtly to 64% of alarms, which we read as a rational adaptation and not a lapse in diligence.
We record four facts per page in real time, namely the trigger, the first thing checked, the first action and if the alert changed the plan. Then we divide action-changing alerts by total alerts over two weeks and track time to first action.
No, in our view. Fewer alerts help only when the ones that remain are actionable and supported. Cutting volume can concentrate hard, unfamiliar problems on the on-call engineer, so we pair volume with actionability and time to first action.
We want a runbook with the first diagnostic commands, a view of upstream and downstream dependencies, and a clear statement of what action the page expects. A page without those is a request for a search, so it should say so.
About the author
Operate Team
The team behind Operate
Operate Team builds Operate, a self-hosted AI SRE that reads your logs, databases and code to find the root cause of production issues with evidence, then drafts the fix as a patch for an engineer to review. Operate runs in your own infrastructure with read-only access to your systems.
