← Back to the framework

From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol

N = 16 Medical imagingClinical assessment / diagnosis

This study in the framework

Human Inherited
Expertise level Manipulated
Experts - juniorExperts - senior
Team composition Individual
Task Inherited
Difficulty Manipulated
LowHigh
Stakes High
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance Not applicable
AI design
Protocol Human-first (update)
AI stance Reflective
Interactivity Static
XAI
Presence Yes
Type Visual / saliency · Contrastive
Quality Not reported
Number of AI advisors One
Research questions

RQ1.1 Does judicial support improve accuracy? RQ1.2 Does judicial support improve accuracy for the most complex cases? RQ2.1 Does judicial support improve confidence in the final decision? RQ2.2 Does judicial support improve confidence in the final decision for the most complex cases? RQ3.1 Is judicial support perceived as useful? RQ3.2 Is judicial support perceived as more useful by less vs more expert X-ray readers?

In the authors' words

"we propose 'Judicial AI,' an innovative interaction protocol aimed at reducing automation bias and preserving a sense of agency. This system presents contrasting explanations to medical professionals rather than definitive recommendations, encouraging user engagement and critical evaluation"

abstract

"our findings show a significant improvement in diagnostic accuracy for complex cases among experienced users (p = .045), with an overall accuracy increase of 0.24 [...] However, the protocol was less beneficial for less experienced users, suggesting that cognitive load might be a limiting factor"

abstract

"we discovered a unique case of automation bias where the user's mental model of the system is less relevant, but the persuasiveness of the explanations is critical, like in the case of the so called white-box paradox"

discussion

"less experienced users rated the judicial support as more fatiguing and less useful. This result may be attributed to the cognitive load imposed by the need to interpret multiple explanations rather than being presented with a definitive answer"

discussion
Experimental design

An exploratory within-subject before-and-after study of a novel interaction protocol the authors call Judicial AI, contrasted conceptually rather than experimentally with conventional oracular AI. The system issues no recommendation at all: for each case it presents two activation maps, one supporting the presence of a fracture and one supporting its absence, and leaves the verdict to the clinician. Every participant experienced the same protocol, so the comparison is each clinician's unaided judgement against their own post-support judgement.

Participants. 16 medical professionals recruited on a convenience basis, 8 specialised spine surgeons and 8 musculoskeletal radiologists. Responses were stratified by expertise, dichotomised at 10 years of experience into less experienced and more experienced.

Task. Detecting traumatic thoraco-lumbar fractures from vertebral X-rays, a task the authors select precisely because it is hard: the false negative rate in practice remains close to 30%. 18 images were chosen by the clinician authors from the held-out test set of a dataset of 630 cropped vertebral X-rays derived from 151 trauma patients at a Milan spine surgery centre, 302 positive and 328 negative, annotated by three board-certified spine surgeons against a CT and MRI gold standard. The 18 were selected as good examples of positive and negative cases of varying diagnostic complexity.

Protocol. A two-page LimeSurvey sequence per case, human-first by construction. Page one showed the high-resolution X-ray and collected the initial fracture judgement, confidence on a six-point semantic differential and a rating of the perceived complexity of the case. Page two showed three thumbnails, the original image and the two contrasting activation maps, any of which could be opened at full resolution, together with a reminder of the initial judgement and controls to confirm or change it, plus final confidence and a rating of the utility of the support. Confidence scores produced by the models were deliberately not disclosed to participants.

Full findings

Withholding the recommendation did not cost accuracy and helped where it should, on hard cases and in expert hands, but the benefit did not reach less experienced clinicians and a fifth of participants were talked out of correct answers by the map arguing the wrong side.

Accuracy. The overall effect was positive but not significant, Glass's Delta 0.24 [-0.10, 0.59], p = .358. It concentrated in two strata. For more experienced clinicians the effect was large and significant, Delta 0.99 [0.50, 1.47], p = .045. For more complex cases the interval excluded zero, Delta 0.35 [0.039, 0.656], though p = .181. In the two complementary strata the point estimate was slightly negative: less complex cases Delta -0.146 [-0.493, 0.149], p = .47, and less expert participants Delta -0.148 [-0.45, 0.161], p = .75.

Non-inferiority is the more important result, since the design question was whether removing the recommendation does harm. Judicial support was non-inferior to unaided diagnosis overall (Z = 3.9425, p < .001) and for more experienced doctors (Z = 4.4117, p < .001), but non-inferiority could not be established for less experienced doctors (Z = 1.2491, p = .1058). The authors flag this as unexpected and as the study's main caution: a protocol that demands interpretation may not be safe for those least equipped to interpret.

Automation bias without an authority to defer to. Four participants did not improve and three, 19% of the sample, got worse by switching initially correct diagnoses to incorrect ones after seeing the maps. The authors emphasise what makes this unusual: the system asserted nothing and offered both sides, so the bias cannot be explained by perceived superiority or authority of the machine. Both maps were coherent with the option they backed, but one of each pair was irrelevant to the task, and participants who deteriorated were swayed by the irrelevant one. They name this a case where the user's mental model matters less than the persuasiveness of the explanation, and connect it to their own white-box paradox.

Confidence rose more reliably than accuracy. Cliff's Delta was 0.191 [0.05, 0.33] overall (p = .058), 0.296 [0.151, 0.442] for more complex cases (p = .034), 0.33 [0.203, 0.451] for more experienced participants (p = .082) and only 0.062 [-0.077, 0.203] for less experienced ones (p = .317). Non-inferiority on confidence held in every stratum including the less experienced (overall W = 1715.0, p < .001; more experienced W = 968.0, p < .001; less experienced W = 2131.0, p = .001). Confidence gains are therefore the one benefit that reached everyone, while accuracy gains did not.

Clinical translation. The Number of Needed Decisions to avoid one misdiagnosis is 42.1, reached in roughly twelve days at the hospital that supplied the data, implying about 26 diagnostic errors avoided per year against roughly 1,080 patients discharged with that diagnosis.