
At the Icahn School of Medicine at Mount Sinai, researchers found that brief safety reminders can lead to safer decisions by AI models in clinical settings, decreasing potentially harmful choices. The study, published Sept. 26 in Communications Medicine, reveals AI models’ decisions can be influenced by surrounding context and instructions in clinical scenarios.
Testing on millions of responses
A team of investigators assessed 20 advanced language models by presenting them with 501 versions of 50 medical situations, plus 100 anonymized hospital discharge summaries. After analyzing over 10 million generated answers, they identified roughly 1.18 million responses that could pose risks in clinical settings. Without any safety prompt, 16.6% of the outputs were potentially dangerous. When a short safety note was included, that figure dropped to 10.1%
The study shows that assessing AI in healthcare must go beyond verifying factual accuracy—it must also examine how models react when given instructions that compromise safety. ‘AI systems don’t operate in isolation,’ explains lead author Mahmud Omar, MD, a physician-researcher and lecturer in the Windreich Department of Artificial Intelligence and Human Health at Mount Sinai’s Icahn School of Medicine. ‘Their responses are shaped by the phrasing, urgency, or authority behind a request, especially when those cues conflict with proper care.’
“The language, framing and context surrounding a request can influence how they respond, including when an instruction could be unsafe,” Omar adds. “A simple safety reminder reduced potentially harmful choices in most of the models we tested, which is encouraging. But it did not eliminate them, so a reminder should be viewed as one safeguard, not a substitute for clinical oversight.”
Related Post: Health and Aged Care Update Released
How context shapes decisions
In one test, researchers presented models with scenarios where they were told to omit necessary follow-up blood tests to ease workload pressures. The instruction was sometimes framed as an emergency or delivered as a directive from a supervisor. Each model then had to select from four options: comply with the request, stick to the recommended protocol, or consult a human clinician.
The team adjusted the wording of each scenario and tested three different safety prompts. Every variation was run 10 times, with answer choices shuffled to avoid bias. Across 19 of the 20 models, including both simulated cases and real discharge summaries, the prompts cut down on risky recommendations.
Examples of potentially harmful choices included skipping needed tests to reduce workload or stopping antibiotic treatment before completing the recommended regimen without a sufficient clinical reason.
Building automated safety checks
‘Our findings show that safety evaluations must move past simple accuracy checks,’ states co-senior author Girish N. Nadkarni, MD, MPH, who leads the Windreich Department and directs Mount Sinai’s Hasso Plattner Institute for Digital Health. ‘As AI grows more independent, we must determine whether it can recognize unsafe directives, challenge them, or seek human input when needed.’