← Back to the framework

AI Assistance in Medical Decision-Making: The Role of Recommendations and Explanations in Simulated Clinical Cases

N = 34 Clinical assessment / diagnosis

This study in the framework

Human Inherited
Expertise level Experts - junior
Team composition Individual
Task Inherited
Difficulty Manipulated
LowMedium
Stakes High
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 95%-100%
AI design
Protocol On request · No AI (control) · AI-first
AI stance Directive
Interactivity Static
XAI Manipulated
Presence Manipulated
Type None · Narrative
Quality Correct
Number of AI advisors One
Research questions
  • RQ1. How do AI recommendations and explanations influence diagnostic accuracy and exam prescription appropriateness throughout the holistic clinical assessment process?
  • RQ2. How do AI recommendations and explanations impact cumulative decision-making efficiency (time and resource use) across sequential clinical stages?
  • RQ3. How do these effects vary with case complexity within a holistic decision-making framework?
In the authors' words

"Results indicate that AI recommendations and explanations can enhance medical performance, but do not consistently improve efficiency."

abstract

"Interestingly, none of the participants in the AI condition used this specific phrasing, whereas 46% of those in the XAI condition did. This indicates a higher degree of reliance on the AI system when explanations were provided alongside recommendations."

discussion, reliance

"We believe that the AI recommendation in this result heightened participants' awareness of the diagnostic challenge, prompting them to adopt a more cautious and deliberative approach. While this helped accuracy, it also led to over-prescription, which runs counter to our hypothesis that AI would improve exam efficiency (H2)."

discussion
Experimental design

Between-subjects on AI assistance level, within-subjects on case complexity, with the order of the two cases counterbalanced.

Task. Two simulated emergency-room clinical cases on UpHill Simulate, a gamified commercial platform for digital case studies, presented in Portuguese and using no real patient data. Each case runs the full clinical sequence: gather patient details, perform a virtual physical examination, prescribe diagnostic exams, read the results, give a primary diagnosis, then set a treatment plan. The platform scores each action against a predefined list of expected actions weighted by clinical importance.

Operationalising complexity. One case was chosen to be harder than the other. PTE, the low-difficulty case, is a female patient with pulmonary thromboembolism whose symptoms point clearly to the diagnosis; the standard workup is electrocardiogram, Angio-CT, chest X-ray and D-dimers. PAC, the medium-difficulty case, is a male patient with sepsis arising from pneumonia, where the difficulty is that the symptoms overlap with a localised infection and distinguishing sepsis requires integrating several cues; the key signal is an arterial blood gas showing lactate at 19 mmol/L against a normal ceiling of 2.5.

Conditions. Control solved both cases unaided. AI received the agent's recommendations. XAI received the same recommendations with textual explanations attached. Assignment was random.

AI system. A scripted agent called Dra. Aida, not a deployed model. Recommendations and explanations were generated in one pass by GPT-4 from patient data, exam results and structured prompts, then curated and checked for content accuracy by an expert clinician. The agent intervenes at two of the three decision points: it suggests three relevant exams from the patient information, and it suggests the most probable diagnosis for each prescribed exam, reasoning from that exam alone without accumulating evidence across exams. Treatment planning was deliberately left unassisted, because fifth-year students have no treatment training and AI support there would have manufactured overreliance rather than simulated practice. Participants had to click a button to open the recommendations, so exposure was self-initiated rather than pushed.

Deliberate imperfections in the advice. In the PTE case the Angio-CT, the exam that actually settles the diagnosis, was removed from the agent's exam recommendations during curation, to see whether students would order it themselves; a chest X-ray, useful only for differential diagnosis, was recommended in its place. In the PAC case the arterial blood gas was recommended and is the only exam from which the agent proposes sepsis; the agent's diagnoses from the other exams point at pneumonia, which is not wrong but misses the condition that needs urgent treatment.

Measures. Behavioural: number of exams prescribed, split into necessary and unnecessary against the platform's expected actions; a normalised Exam Score and a normalised Clinical Case Score, both weighting actions by clinical importance; reliance at the exam stage as the share of AI-suggested exams the participant ordered, plus tracking of the two diagnostic exams named above; correct diagnosis as a binary against a clinician-verified list of acceptable ICD-10-CM equivalents; exact diagnosis keyword as a binary for reproducing the agent's exact wording, used as the diagnosis-stage reliance measure; and case resolution time. Non-verbal: eye tracking with a Tobii Pro Nano at 60 Hz, I-VT fixation classification, reporting fixation count and total fixation duration; and mouse tracking of clicks on recommendations and time spent reading them.

Full findings

AI assistance improved what students decided and not how efficiently they decided it, and the two effects separate by case complexity: the easy case got faster, the hard case got more accurate and more wasteful.

On accuracy, the Clinical Case Score, which combines exam prescription and diagnosis, was higher with AI assistance in both cases. In PTE the three-way test missed conventional significance (F(2,31) = 3.108, p = .059) and the pooled contrast of Control against AI + XAI was reported as significant (p = .02). In PAC the three-way test was significant (F(2,31) = 4.517, p = .019) with XAI outperforming Control (p = .015). The clearest single result in the paper is diagnostic: 63.6% of the XAI group identified sepsis against 41.7% in AI and 18.2% in Control, and although that comparison itself fell short (p = .086), the matching exact-keyword measure was significant (p = .007), with 45.5% of XAI participants using the agent's exact phrase "Severe Sepsis Without Septic Shock" against 9.1% in Control and 0.0% in AI. Explanations, not recommendations alone, moved the hard diagnosis. H1.1 is supported; H1.2 only partially, since explanations did nothing for exam prescription.

Efficiency went the wrong way in the hard case. Students with AI assistance ordered more exams in PAC (8.42 in AI and 8.27 in XAI against 5.91 in Control) and more unnecessary ones (4.75 and 4.91 against 2.55), neither difference significant at this sample size. Resolution time in PAC did not differ across conditions (F(2,31) = .803, p = .457). H2 is contradicted and H3 holds only for the easy case, where AI assistance cut resolution time from 20.07 minutes in Control to 15.34 in AI and 13.97 in XAI (F(2,31) = 3.350, p = .048, the post hoc difference falling between Control and XAI). Eye tracking agrees: total fixation duration in PTE dropped from 1013.9 s in Control to 629.8 s in AI and 680.7 s in XAI (F(2,27) = 5.842, p = .008), with no such effect in PAC. The authors read the shorter fixations as reduced attentional demand, and the pattern as AI turning an easy case into confirmation of a hypothesis the student already held.

Reliance rose with explanations, and the mouse-tracking data is the strongest evidence for it. In PAC, XAI participants clicked through to the recommendations more often than AI participants (F(2,21) = 10.470, p = .004) and spent far longer reading them (14.089 s against 4.406 s, p < .001); the reading-time gap held in PTE too (14.127 s against 6.047 s, p = .002). The behavioural reliance measures point the same way without reaching significance: reliance on suggested exams in PTE ran 0.509 in XAI, 0.417 in AI and 0.309 in Control, and the optional chest X-ray, which the agent recommended but which is redundant once an Angio-CT is ordered, was prescribed by 72.7% of XAI and 66.7% of AI participants against 27.3% of Control (p = .061). That last figure is the paper's clearest behavioural trace of following the agent past the point of usefulness, and the authors present it as reliance rather than as overreliance.

The authors are explicit about why they cannot call it overreliance: the design contains no incorrect AI recommendations, so appropriate and inappropriate reliance cannot be separated. This is the study's main structural limit, and it is stated rather than glossed. The two deliberate imperfections that were built in, dropping the Angio-CT from the PTE recommendations, and having the agent read pneumonia rather than sepsis from every exam but one, are omissions and partial readings rather than errors, and neither is analysed as an error condition.