From Oracular to Judicial: Enhancing Clinical Decision Making through Contrasting Explanations and a Novel Interaction Protocol
N = 16 Medical imagingClinical assessment / diagnosisThis study in the framework
Reported, not manipulated, and unusually beside the point: the system issues no prediction to be accurate about. Three ResNeXt-50 models transferred from ImageNet to binary fracture classification, one optimised for accuracy, one for true positive rate and one for true negative rate, with hyperparameters tuned on a validation set using Optuna. The dataset was 630 cropped vertebral X-rays from 151 trauma patients, 302 positive and 328 negative, labelled by three board-certified spine surgeons against a CT and MRI gold standard, split 80/10/10 with standard augmentation.
RQ1.1 Does judicial support improve accuracy? RQ1.2 Does judicial support improve accuracy for the most complex cases? RQ2.1 Does judicial support improve confidence in the final decision? RQ2.2 Does judicial support improve confidence in the final decision for the most complex cases? RQ3.1 Is judicial support perceived as useful? RQ3.2 Is judicial support perceived as more useful by less vs more expert X-ray readers?
"we propose 'Judicial AI,' an innovative interaction protocol aimed at reducing automation bias and preserving a sense of agency. This system presents contrasting explanations to medical professionals rather than definitive recommendations, encouraging user engagement and critical evaluation"
abstract
"our findings show a significant improvement in diagnostic accuracy for complex cases among experienced users (p = .045), with an overall accuracy increase of 0.24 [...] However, the protocol was less beneficial for less experienced users, suggesting that cognitive load might be a limiting factor"
abstract
"we discovered a unique case of automation bias where the user's mental model of the system is less relevant, but the persuasiveness of the explanations is critical, like in the case of the so called white-box paradox"
discussion
"less experienced users rated the judicial support as more fatiguing and less useful. This result may be attributed to the cognitive load imposed by the need to interpret multiple explanations rather than being presented with a definitive answer"
discussion
Experimental design
An exploratory within-subject before-and-after study of a novel interaction protocol the authors call Judicial AI, contrasted conceptually rather than experimentally with conventional oracular AI. The system issues no recommendation at all: for each case it presents two activation maps, one supporting the presence of a fracture and one supporting its absence, and leaves the verdict to the clinician. Every participant experienced the same protocol, so the comparison is each clinician's unaided judgement against their own post-support judgement.
Participants. 16 medical professionals recruited on a convenience basis, 8 specialised spine surgeons and 8 musculoskeletal radiologists. Responses were stratified by expertise, dichotomised at 10 years of experience into less experienced and more experienced.
Task. Detecting traumatic thoraco-lumbar fractures from vertebral X-rays, a task the authors select precisely because it is hard: the false negative rate in practice remains close to 30%. 18 images were chosen by the clinician authors from the held-out test set of a dataset of 630 cropped vertebral X-rays derived from 151 trauma patients at a Milan spine surgery centre, 302 positive and 328 negative, annotated by three board-certified spine surgeons against a CT and MRI gold standard. The 18 were selected as good examples of positive and negative cases of varying diagnostic complexity.
Protocol. A two-page LimeSurvey sequence per case, human-first by construction. Page one showed the high-resolution X-ray and collected the initial fracture judgement, confidence on a six-point semantic differential and a rating of the perceived complexity of the case. Page two showed three thumbnails, the original image and the two contrasting activation maps, any of which could be opened at full resolution, together with a reminder of the initial judgement and controls to confirm or change it, plus final confidence and a rating of the utility of the support. Confidence scores produced by the models were deliberately not disclosed to participants.
Full findings
Withholding the recommendation did not cost accuracy and helped where it should, on hard cases and in expert hands, but the benefit did not reach less experienced clinicians and a fifth of participants were talked out of correct answers by the map arguing the wrong side.
Accuracy. The overall effect was positive but not significant, Glass's Delta 0.24 [-0.10, 0.59], p = .358. It concentrated in two strata. For more experienced clinicians the effect was large and significant, Delta 0.99 [0.50, 1.47], p = .045. For more complex cases the interval excluded zero, Delta 0.35 [0.039, 0.656], though p = .181. In the two complementary strata the point estimate was slightly negative: less complex cases Delta -0.146 [-0.493, 0.149], p = .47, and less expert participants Delta -0.148 [-0.45, 0.161], p = .75.
Non-inferiority is the more important result, since the design question was whether removing the recommendation does harm. Judicial support was non-inferior to unaided diagnosis overall (Z = 3.9425, p < .001) and for more experienced doctors (Z = 4.4117, p < .001), but non-inferiority could not be established for less experienced doctors (Z = 1.2491, p = .1058). The authors flag this as unexpected and as the study's main caution: a protocol that demands interpretation may not be safe for those least equipped to interpret.
Automation bias without an authority to defer to. Four participants did not improve and three, 19% of the sample, got worse by switching initially correct diagnoses to incorrect ones after seeing the maps. The authors emphasise what makes this unusual: the system asserted nothing and offered both sides, so the bias cannot be explained by perceived superiority or authority of the machine. Both maps were coherent with the option they backed, but one of each pair was irrelevant to the task, and participants who deteriorated were swayed by the irrelevant one. They name this a case where the user's mental model matters less than the persuasiveness of the explanation, and connect it to their own white-box paradox.
Confidence rose more reliably than accuracy. Cliff's Delta was 0.191 [0.05, 0.33] overall (p = .058), 0.296 [0.151, 0.442] for more complex cases (p = .034), 0.33 [0.203, 0.451] for more experienced participants (p = .082) and only 0.062 [-0.077, 0.203] for less experienced ones (p = .317). Non-inferiority on confidence held in every stratum including the less experienced (overall W = 1715.0, p < .001; more experienced W = 968.0, p < .001; less experienced W = 2131.0, p = .001). Confidence gains are therefore the one benefit that reached everyone, while accuracy gains did not.
Clinical translation. The Number of Needed Decisions to avoid one misdiagnosis is 42.1, reached in roughly twelve days at the hospital that supplied the data, implying about 26 diagnostic errors avoided per year against roughly 1,080 patients discharged with that diagnosis.