Case Study

A clinical alert dashboard, designed to cut hospitalist override

Physicians override 95% of drug-interaction alerts, across 18,354 medication orders at two hospitals, with no improvement since the Meaningful Use era began. (Bryant et al., Applied Clinical Informatics, 2014)

Role

Solo: research, synthesis, design, testing, prototyping

Timeline

May to Aug 2026

Tools

Figma, Stark, Useberry, Claude

100%
Correctly triaged the critical alert alone
73.3
SUS score, Round 1 (68+ is above average)
33%
Escalated it once AI confidence was also stated, up from 0% in Round 1
ESI-1 alert expanded: AI confidence, evidence basis, and the Hold/Override/Escalate action row

The ESI-1 alert, expanded. Round 1 design: Hold is styled primary; nothing states the AI's recommendation.

ESI-1 alert expanded in the Round 2 design: Escalate to Pharmacy styled as the primary button, with a stated recommendation banner reading 94% confidence supports escalation to pharmacy

The Round 2 fix: Escalate restyled as primary, plus a stated recommendation above the button row. Testing showed this wasn't enough on its own. See Testing below for why.

30-second summary

Problem

Every drug-interaction alert looks identical on screen, whether it's routine or life-threatening. Hospitalists override 95% of them, not carelessness, but alarm fatigue: so many alerts fire that people stop reacting to any of them. That's tied to real patient harm.

Approach

I didn't have access to real hospitalists, so I built the problem from 18 clinical sources, then tested the design across two rounds with nurses as the closest available stand-in.

Key decision

A triage table using ESI, the severity scale hospitals already use, with the AI's confidence shown right next to the action, and friction that scales with risk: one click to accept, a short reasoned form to override.

Status

Two rounds of testing are done. Every tester found the critical alert right away, in both rounds. Getting people to act on the AI's recommendation improved but didn't fully work: nobody did in round one, and moving the recommendation next to the button got one of three testers to act on it in round two. The other two missed it for two different reasons, which is the real finding.

Problem

Hospitalists handle the highest volume of daily medication orders in inpatient care, and drug-interaction alerts are supposed to catch what they miss. But every alert renders at identical visual priority: no position, size, or color signal separates a life-threatening interaction from a routine one. Volume alone isn't the failure; making everything look equally urgent trains clinicians to ignore all of it.

39-fold

antibiotic overdose in a 2015 UCSF case, after a physician and pharmacist both overrode a critical alert that "looked exactly the same" as routine ones.

Wachter 2015

216+ deaths

in the U.S. (2005–2010) tied to alarm fatigue, per a Boston Globe investigation.

Wachter 2015

187/day

alerts per patient at UCSF ICUs, almost all clinically insignificant, 2,507,822 unique alarms across five ICUs in one month.

Drew et al. 2014

+13% risk

per interruption during medication administration. Four interruptions doubles the rate of errors likely to cause permanent harm or death.

Westbrook et al. 2010

The Joint Commission issued an urgent alarm-safety directive in 2013 (updated 2016), naming alarm-related problems a top health-technology hazard. (The Joint Commission 2013/2016)

Solution

Three things the design needed to answer: the minimum info a hospitalist needs to triage an alert without leaving the alert view, what signal reads as "critical" before any text gets read, and when an AI confidence score actually changes behavior instead of getting overridden. The answer: a triage table with ESI severity tiers already in clinical vocabulary, AI confidence surfaced inline, and friction that scales with risk: one click to accept the recommendation, a reasoned form to override it.

A table, color as backup only

Every alert is a row: severity, drug ordered vs. drug on file, AI confidence, recommended action, all visible without expanding. Strip the color out and the ESI-1 row is still identifiable by position, height, and weight; color-only encoding fails color-deficient clinicians and is unreliable under time pressure regardless.

Mazur et al. 2019; NN/g 2018; The Joint Commission 2013/2016

ESI, not an invented scale

The panel borrows the Emergency Severity Index, already in a hospitalist's vocabulary, instead of a made-up "High/Medium/Low." "ESI-1" needs zero explanation mid-shift.

Hussain et al. 2019

Friction scales with risk

The recommended action (Hold) completes in one click. Overriding requires picking a reason from five visible radio options, not a dropdown, before Confirm enables. Dropdown reasons are too coarse; clinicians pick the fastest option, not the accurate one.

Payne et al. 2015

Cut override at the point of highest risk

Override drops from 99.3% to 1.7% once a high-confidence recommendation is actually stated, not left sitting in a separate evidence panel, across 6,689 cardiovascular cases. That's the lever for the numbers above: Round 1 styled Escalate as secondary with nothing stated; Round 2 made it primary and added a one-line stated recommendation.

Override rate, by AI confidence band

90–99%
1.7%
70–79%
99.3%

Yu et al. 2025; NN/g 2025; NN/g 2018; Payne et al. 2015

Process

1. Research

Built the problem from 18 peer-reviewed and clinical sources instead of primary interviews. Synthesis produced two personas and the ESI-based severity framing the design is built around.

2. Ideation in Stitch

Used Google's Stitch tool to quickly generate and compare three distinct dashboard directions, then narrowed down to the strongest one before committing time to a full Figma build.

3. Wireframes

Sketched three low-fidelity layout directions and ran each through a 3-second test: can a clinician spot the most critical alert without reading anything? A compact triage table passed as the best direction over a single-column card where height signaled priority over severity index.

4. Figma: refinement and prototyping

Built the chosen direction into a full interactive prototype: the triage table, expandable alert rows, the explainability panel, and the action buttons. Ran it through a WCAG 2.2 AA accessibility check in Stark before testing.

5. Testing, round 1

Sent the prototype to three nurse proxy testers for an unmoderated usability test: severity understanding of the alert triage, an escalate-or-override decision, and an override-friction task, plus two standard surveys (SUS, NASA-TLX).

6. Revision

Every tester skipped the AI's recommended action on the highest-confidence alert. Restyled Escalate as the primary button and added a one-line stated recommendation above it, instead of leaving the connection between confidence score and action implicit.

7. Testing, round 2

Sent the revised prototype to the same three testers. One acted on the AI's recommendation, the other two still didn't, each for a different reason. The numbers behind both rounds are below.

Round 1 results (3 testers)

73.3
Usability score (SUS), 68+ = above avg
35.0
Mental effort score (NASA-TLX), lower is better
100%
Task 1: found the critical alert
0%
Task 2: acted on a 94%-confidence AI recommendation
67%
Task 3: finished the override form

All three testers found the critical alert on their own, in about two minutes on average. None of them acted on the AI's recommendation once it was on screen: two chose to override it, one chose to just hold. Two of three finished the override form; the third dropped off partway through, and a tester dropped off at that exact same step again in round two.

Round 1 (3 testers) → Round 2 (3 testers)

SUS Score

R1
73.3
R2
74.2

AI-confidence decision

R1
0%
R2
33%

Finished the override form

R1
67%
R2
67%

Round 2 shows the fix partly worked: one of three testers acted on the AI's recommendation, up from zero in round one. The usability score held steady, 73.3 to 74.2. But the three testers got to their answers in very different ways. One looked at the screen for 11 seconds and made a single click before choosing Hold, likely too fast to have registered the new stated recommendation. Another spent 87 seconds and made three clicks before choosing Override anyway, long enough to have read it and rejected it. The third spent two minutes, clicked directly on "Escalate to Pharmacy," and confirmed it a second later. Three different outcomes from the same fix: didn't see it, saw it and rejected it, saw it and acted on it. The first two are different problems, and a bigger button could only fix one of them. The override-form drop-off from round one also happened again, with a tester hitting the same dead end a second time; that's now a confirmed problem.

What I learned

Most problems are not strictly visual

Stating the AI's recommendation and styling it as primary got one of three Round 2 testers to act on it, up from zero in Round 1. The other two picked something else, for different reasons: one decided in 11 seconds, too fast to have read the new text at all; the other spent 87 seconds and rejected the recommendation anyway. One is a visibility problem. The other is a trust problem: understanding how different users react to and decide on critical information is a complex matter that can't be solved with a single hierarchy change.

Recruiting is its own project risk

Getting qualified testers took longer than I accounted for. Access to clinicians willing to test an unpaid prototype is genuinely limited, that's a project risk that even when planned for, can have unexpected pitfalls such as platform bugs, panel narrowing, or just delay in response times.

Where this stands

Two rounds in, the override problem this project set out to fix moved but didn't close: escalating the highest-risk alert went from 0% to 33% once the recommendation was stated and made primary. That's real progress against the fatigue-driven override pattern, but two of three testers still didn't act on it, and recruiting for this panel has already gone over allotted time. Running a full third round isn't in my scope, but if I were to build on this further, I'd move to a moderated session so I can talk directly with each tester about why they overrode or held instead of escalating, before moving to deeper research analysis or design refinement.

Sources

Every citation opens in a new tab.

  1. Bryant AD, Fletcher GS, Payne TH. "Drug interaction alert override rates in the Meaningful Use era: no evidence of progress." Applied Clinical Informatics, 2014. ↗
  2. Wachter R. "Beware of the Robot Pharmacist" (the Pablo Garcia Septra overdose case). Medium, 2015. ↗
  3. The Joint Commission. Sentinel Event Alert Issue 50: Medical device alarm safety in hospitals. 2013, updated 2016. ↗
  4. Drew BJ, Harris P, Zègre-Hemsey JK, et al. "Insights into the problem of alarm fatigue with physiologic monitor devices" (UCSF ICU observational study). PLoS One, 2014. ↗
  5. Westbrook JI, Woods A, Rob MI, et al. "Association of interruptions with an increased risk and severity of medication administration errors." Archives of Internal Medicine, 2010. ↗
  6. Payne TH, Hines LE, Chan RC, et al. "Recommendations to improve the usability of drug-drug interaction clinical decision support alerts." Journal of the American Medical Informatics Association, 2015. ↗
  7. Hussain MI, Reynolds TL, Zheng K. "Medication safety alert fatigue may be reduced via interaction design and clinical role tailoring." JAMIA, 2019. ↗
  8. Mazur LM, Mosaly PR, Moore C, Marks LB. "Association of the Usability of Electronic Health Records With Cognitive Workload and Performance Levels Among Physicians." JAMA Network Open, 2019. ↗
  9. Yu Y, Gomez-Cabello CA, Haider SA, et al. "Enhancing Clinician Trust in AI Diagnostics: A Dynamic Framework for Confidence Calibration and Transparency." Diagnostics, 2025. ↗
  10. Nielsen Norman Group. "Dashboards: Making Charts and Graphs Easier to Understand." 2018. ↗
  11. Sunwall E. "Prioritize Smarts over Sentience to Increase Trust with AI." Nielsen Norman Group, 2025. ↗