← Back to the framework

The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems

N = 7 Clinical assessment / diagnosis

This study in the framework

Human Inherited
Expertise level Experts - unspecified
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance <60%
AI design
Protocol AI-first · Confidence display
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Feature importance
Quality Correct
Number of AI advisors One
Research questions
  • RQ1. What effects do confidence explanations have on users trusting a CDSS and relying on system suggestions?
  • RQ2. What effects do why explanations have on users trusting a CDSS and relying on system suggestions?
  • RQ3. What should be explained to help clinicians better assess the appropriateness of the system's suggestion?
In the authors' words

"Our results show that the amount of system confidence had only a slight effect on trust and reliance. More importantly, giving a fuller explanation of the facts used in making a diagnosis had a positive effect on trust but also led to over-reliance issues, whereas less detailed explanations made participants question the system's reliability and led to self-reliance problems."

abstract

"CDSS users who trust the system highly are also likely to over-rely on the system's suggestions, while users who distrust the system are likely to rely on their own knowledge, even if it is poor."

discussion

"Whilst a more detailed explanation may promote over-reliance, we argue that providing no explanation at all is not a viable option as they are desirable and necessary."

discussion
Experimental design

Exploratory between-groups study. Seven healthcare practitioners diagnosed eight clinical vignettes describing fictional patients with balance-related disorders, using a Wizard-of-Oz clinical decision support prototype. Participants entered medical history, symptoms and examination results, received a suggested diagnosis, and either accepted or rejected it while thinking aloud.

Two factors were varied. Between groups, the why explanation listed either all items of medical history, symptoms and examination results (comprehensive version, four participants) or examination results only (selective version, three participants). Within participants, each suggestion carried a confidence percentage set either above 75% or below 30%, balanced across vignettes. Four of the eight suggested diagnoses were incorrect by design, and incorrect suggestions were built to share symptoms or examination results with the correct diagnosis. Fifty-two vignettes were completed in total, twenty-eight in the comprehensive group and twenty-four in the selective group.

Trust was rated on a seven-point scale before and after use. Reliance was derived from agreement with system suggestions and from whether the resulting decision was right or wrong. No statistical tests were run; analysis was qualitative, combining raw counts with thematic coding of think-aloud and interview data.

Full findings

Confidence percentage had little effect on reliance. Participants agreed with 21 high-confidence suggestions and 18 low-confidence suggestions out of 52 shown, split equally between the two levels, and only one participant cited system confidence as a reason to trust the system. Four participants stated that they did not understand what the percentage represented.

Explanation detail did not change overall accuracy: the comprehensive group made 16 right decisions and the selective group 14. It did change the direction of error. The comprehensive group agreed with 11 incorrect suggestions against 7 in the selective group, and disagreed with only one correct suggestion against three, indicating over-reliance under comprehensive explanations and self-reliance under selective ones. Three of the four comprehensive participants raised their trust rating after use, against one of three in the selective group.

Participants attributed the persuasiveness of comprehensive explanations to three impressions: that the system drew on up-to-date medical knowledge, that it identified the salient features of a case, and that it reasoned as they would. Selective explanations were read as evidence that the system ignored history and symptoms, and so applied a reasoning process inferior to their own.

Participants requested four further kinds of explanatory information: what the confidence figure means and how it was derived, a description of the disorder and of a typical case against which to check the fit, the pathological link and the feature weights behind the suggestion, and differential diagnoses.