Case Study
Physicians override 95% of drug-interaction alerts, across 18,354 medication orders at two hospitals, with no improvement since the Meaningful Use era began. (Bryant et al., Applied Clinical Informatics, 2014)
The ESI-1 alert, expanded. Round 1 design: Hold is styled primary; nothing states the AI's recommendation.
The Round 2 fix: Escalate restyled as primary, plus a stated recommendation above the button row. Testing showed this wasn't enough on its own. See Testing below for why.
30-second summary
Problem
Every drug-interaction alert looks identical on screen, whether it's routine or life-threatening. Hospitalists override 95% of them, not carelessness, but alarm fatigue: so many alerts fire that people stop reacting to any of them. That's tied to real patient harm.
Approach
I didn't have access to real hospitalists, so I built the problem from 18 clinical sources, then tested the design across two rounds with nurses as the closest available stand-in.
Key decision
A triage table using ESI, the severity scale hospitals already use, with the AI's confidence shown right next to the action, and friction that scales with risk: one click to accept, a short reasoned form to override.
Status
Two rounds of testing are done. Every tester found the critical alert right away, in both rounds. Getting people to act on the AI's recommendation improved but didn't fully work: nobody did in round one, and moving the recommendation next to the button got one of three testers to act on it in round two. The other two missed it for two different reasons, which is the real finding.
Hospitalists handle the highest volume of daily medication orders in inpatient care, and drug-interaction alerts are supposed to catch what they miss. But every alert renders at identical visual priority: no position, size, or color signal separates a life-threatening interaction from a routine one. Volume alone isn't the failure; making everything look equally urgent trains clinicians to ignore all of it.
39-fold
antibiotic overdose in a 2015 UCSF case, after a physician and pharmacist both overrode a critical alert that "looked exactly the same" as routine ones.
216+ deaths
in the U.S. (2005–2010) tied to alarm fatigue, per a Boston Globe investigation.
187/day
alerts per patient at UCSF ICUs, almost all clinically insignificant, 2,507,822 unique alarms across five ICUs in one month.
+13% risk
per interruption during medication administration. Four interruptions doubles the rate of errors likely to cause permanent harm or death.
The Joint Commission issued an urgent alarm-safety directive in 2013 (updated 2016), naming alarm-related problems a top health-technology hazard. (The Joint Commission 2013/2016)
Three things the design needed to answer: the minimum info a hospitalist needs to triage an alert without leaving the alert view, what signal reads as "critical" before any text gets read, and when an AI confidence score actually changes behavior instead of getting overridden. The answer: a triage table with ESI severity tiers already in clinical vocabulary, AI confidence surfaced inline, and friction that scales with risk: one click to accept the recommendation, a reasoned form to override it.
Every alert is a row: severity, drug ordered vs. drug on file, AI confidence, recommended action, all visible without expanding. Strip the color out and the ESI-1 row is still identifiable by position, height, and weight; color-only encoding fails color-deficient clinicians and is unreliable under time pressure regardless.
Mazur et al. 2019; NN/g 2018; The Joint Commission 2013/2016
The panel borrows the Emergency Severity Index, already in a hospitalist's vocabulary, instead of a made-up "High/Medium/Low." "ESI-1" needs zero explanation mid-shift.
The recommended action (Hold) completes in one click. Overriding requires picking a reason from five visible radio options, not a dropdown, before Confirm enables. Dropdown reasons are too coarse; clinicians pick the fastest option, not the accurate one.
Override drops from 99.3% to 1.7% once a high-confidence recommendation is actually stated, not left sitting in a separate evidence panel, across 6,689 cardiovascular cases. That's the lever for the numbers above: Round 1 styled Escalate as secondary with nothing stated; Round 2 made it primary and added a one-line stated recommendation.
Built the problem from 18 peer-reviewed and clinical sources instead of primary interviews. Synthesis produced two personas and the ESI-based severity framing the design is built around.
Used Google's Stitch tool to quickly generate and compare three distinct dashboard directions, then narrowed down to the strongest one before committing time to a full Figma build.
Sketched three low-fidelity layout directions and ran each through a 3-second test: can a clinician spot the most critical alert without reading anything? A compact triage table passed as the best direction over a single-column card where height signaled priority over severity index.
Built the chosen direction into a full interactive prototype: the triage table, expandable alert rows, the explainability panel, and the action buttons. Ran it through a WCAG 2.2 AA accessibility check in Stark before testing.
Sent the prototype to three nurse proxy testers for an unmoderated usability test: severity understanding of the alert triage, an escalate-or-override decision, and an override-friction task, plus two standard surveys (SUS, NASA-TLX).
Every tester skipped the AI's recommended action on the highest-confidence alert. Restyled Escalate as the primary button and added a one-line stated recommendation above it, instead of leaving the connection between confidence score and action implicit.
Sent the revised prototype to the same three testers. One acted on the AI's recommendation, the other two still didn't, each for a different reason. The numbers behind both rounds are below.
Round 1 results (3 testers)
All three testers found the critical alert on their own, in about two minutes on average. None of them acted on the AI's recommendation once it was on screen: two chose to override it, one chose to just hold. Two of three finished the override form; the third dropped off partway through, and a tester dropped off at that exact same step again in round two.
Round 1 (3 testers) → Round 2 (3 testers)
Round 2 shows the fix partly worked: one of three testers acted on the AI's recommendation, up from zero in round one. The usability score held steady, 73.3 to 74.2. But the three testers got to their answers in very different ways. One looked at the screen for 11 seconds and made a single click before choosing Hold, likely too fast to have registered the new stated recommendation. Another spent 87 seconds and made three clicks before choosing Override anyway, long enough to have read it and rejected it. The third spent two minutes, clicked directly on "Escalate to Pharmacy," and confirmed it a second later. Three different outcomes from the same fix: didn't see it, saw it and rejected it, saw it and acted on it. The first two are different problems, and a bigger button could only fix one of them. The override-form drop-off from round one also happened again, with a tester hitting the same dead end a second time; that's now a confirmed problem.
What I learned
Stating the AI's recommendation and styling it as primary got one of three Round 2 testers to act on it, up from zero in Round 1. The other two picked something else, for different reasons: one decided in 11 seconds, too fast to have read the new text at all; the other spent 87 seconds and rejected the recommendation anyway. One is a visibility problem. The other is a trust problem: understanding how different users react to and decide on critical information is a complex matter that can't be solved with a single hierarchy change.
Getting qualified testers took longer than I accounted for. Access to clinicians willing to test an unpaid prototype is genuinely limited, that's a project risk that even when planned for, can have unexpected pitfalls such as platform bugs, panel narrowing, or just delay in response times.
Two rounds in, the override problem this project set out to fix moved but didn't close: escalating the highest-risk alert went from 0% to 33% once the recommendation was stated and made primary. That's real progress against the fatigue-driven override pattern, but two of three testers still didn't act on it, and recruiting for this panel has already gone over allotted time. Running a full third round isn't in my scope, but if I were to build on this further, I'd move to a moderated session so I can talk directly with each tester about why they overrode or held instead of escalating, before moving to deeper research analysis or design refinement.
Every citation opens in a new tab.