← Back to the framework

Understanding the impact of explanations on advice-taking: a user study for AI-based clinical Decision Support Systems

N = 28 Clinical prognosis

This study in the framework

Human Inherited
Expertise level Experts - unspecified
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes High
Stress level / time constraint Not reported
Task uncertainty High (prognostic)
AI Ecosystem Designable
AI performance 95%-100%
AI design
Protocol Human-first (update)
AI stance Directive
Interactivity Interactive/Dynamic
XAI Manipulated
Presence Manipulated
Type None · Feature importance · Counterfactual
Quality Correct
Number of AI advisors One
Research questions
  • RQ1: How do AI explanations impact users' trust in algorithmic recommendations in healthcare?
  • RQ2: How do AI explanations impact users' behavioral intention of using the system in the healthcare context?
In the authors' words

"Our results indicate a more significant impact of advice when an explanation for the DSS decision is provided."

abstract

"We found that participants were keener on taking advice from the AI interface that explained its suggestion than the one that did not."

discussion

"It is interesting to notice that, despite the low perceived explanation quality, participants were influenced by it and relied more on the advice of the AI system."

discussion

"This finding might be in line with previous research on automation bias in medicine, i.e., the tendency to over-rely on automation [34, 40, 49], and will definitely be the subject of future works."

discussion

"Indeed, many participants showed some degree of algorithm aversion and expressed the fear of being replaced by the AI system."

discussion
Experimental design

Within-subjects, two-condition design, with the order of the two interfaces randomised across participants.

Task. Participants estimated, on a 0 to 100% scale, the chance that a patient would suffer an acute myocardial infarction in the near future given the patient's past clinical history. The history was shown as a timeline of visits, each visit a set of grey dots standing for the conditions diagnosed at that visit, which participants could explore by moving the cursor over them. Each participant judged two analogous patient cases, one per interface, the cases chosen to be comparable so that neither condition benefited from learning.

Procedure. For each case the participant gave an initial estimate and rated their confidence in it, then saw the algorithmic suggestion, then gave a final estimate and rated confidence again.

Conditions. Dr.AI presented the algorithmic suggestion alone. Dr.XAI presented the same suggestion together with a domain-aware, removal-based explanation: conditions the algorithm judged relevant to its decision were coloured blue, irrelevant ones stayed grey, and conditions that were absent from the history but would have changed the algorithmic suggestion were shown in yellow. A written summary of the explanation appeared under the suggestion, and selecting a sentence in that summary highlighted the corresponding conditions in the clinical history. Presence of the explanation was the only difference between conditions.

AI system. Doctor AI, a recurrent neural network that predicts a patient's future diagnoses from past clinical histories, adapted from multi-label prediction to binary acute-MI classification. The explanation method is removal-based but deliberately not model-agnostic in the LIME or SHAP sense: features are summarised with regard to their medical meaning, which the authors describe as knowledge-aware.

Measures. Weight of advice, WOA = |F − I| / |A − I|, where I and F are the participant's initial and final estimates and A the algorithmic suggestion, which equalled 100 on every trial. Confidence shift between the pre-advice and post-advice ratings. Explicit trust on a five-point Likert trust scale, behavioural intention on a UTAUT questionnaire, and explanation satisfaction on a five-point explanation satisfaction scale. Open-ended responses were collected and analysed thematically.

Full findings

The explanation moved behaviour and left stated attitudes untouched, and it did so even though participants thought the explanation was poor.

Weight of advice was higher with Dr.XAI (Mdn = 0.31) than with Dr.AI (Mdn = 0), a paired two-sided Wilcoxon signed-rank test giving T = 32.5, p = 0.002, which supports Hp1. The median of 0 in the no-explanation condition is the substantive part of the result: the typical participant did not move their estimate at all when handed a bare algorithmic suggestion, and only moved once the suggestion carried a rationale.

All three attitudinal hypotheses returned nulls. Confidence did not differ between interfaces (T = 169, p = 0.869), contradicting Hp2. Behavioural intention did not differ (T = 37, p = 0.076), contradicting Hp3. Explicit trust did not differ (T = 157.0, p = 0.881), contradicting Hp4. The behavioural effect therefore did not pass through anything participants reported about the system.

The mechanism the authors offer is automation bias rather than persuasion. Participants rated explanation satisfaction low, complained in the open responses of information overload, and relied on the advice anyway. What explanation satisfaction did predict was attitude, not behaviour: it correlated strongly with explicit trust (rs(27) = 0.77, p < .001) and with behavioural intention (rs(27) = 0.67, p < .001). Read together, the presence of an explanation tracks advice-taking while the perceived quality of an explanation tracks stated trust, and the two are separable.

Demand for explanations was high even where they were absent: 54% of participants asked for an explanation in the Dr.AI condition, and 46% expressed no opinion. The qualitative analysis also surfaced a theme the design did not target, algorithm aversion expressed as fear of professional replacement by the system, which the authors argue is an adoption barrier that HCI work on clinical decision support underweights.

One interpretive limit is structural and the authors state it. Only cases the algorithm had classified correctly were used, and the suggestion was always 100, so any increase in weight of advice is by construction movement toward the correct answer. The study can show that explanations increase advice-taking; it cannot separate appropriate from inappropriate reliance, and it does not measure decision accuracy as an outcome. The authors chose this to keep algorithmic accuracy out of the design, and their reading of the WOA result as advice-taking rather than as improved decision quality is careful on this point.