Never tell me the odds: Investigating pro-hoc explanations in medical decision making
N = 16 Medical imagingClinical assessment / diagnosisThis study in the framework
Not applicable. The system makes no prediction, offers no categorical advice and displays no probability score. It abstains from interpreting the case in hand and instead retrieves three precedents from the repository of available cases, each carrying its verified ground-truth diagnosis: the two most similar cases sharing the label the physician has just chosen, and the single most similar case carrying the opposite label.
RQ: What was the impact of showing similar cases instead of regular classifications (the pro-hoc approach) upon decision performance?
"physicians, particularly those with less experience, perceived pro-hoc XAI support as significantly beneficial, even though it did not notably enhance their diagnostic accuracy"
abstract
"the number of decision changes from an initially wrong to a correct diagnosis (11) was more than double the number of decision changes in the reverse direction (5)"
results
"the higher the perceived complexity, the lower the confidence and, most notably, the higher the perceived usefulness of the AI support"
results
Experimental design
A single-condition, within-subjects user study with no control group. Sixteen physicians each judged 18 cases, giving 288 diagnoses.
Participants. Sixteen physicians with varying experience in reading spine x-rays in daily practice: ten board-certified orthopaedic spine subspecialists and six orthopaedic residents.
Task. Identify the presence or absence of vertebral fractures and lesions in spine x-rays presented at 800 by 800 pixels. Eighteen cases, selected in a previous study by the same group for their representativeness of varied and complex cases, balanced between images showing a fracture and images without one. Administered case by case through an online LimeSurvey questionnaire.
Protocol. Each case ran in three steps. First the physician gave an initial binary diagnosis, recorded as HD1, together with the perceived difficulty or complexity of the case and their confidence in the diagnosis, both on six-value ordinal scales. Second, conditional on that initial answer, the system retrieved three cases from the repository of available cases: the two most similar cases carrying the same label the physician had just given, and the single most similar case carrying the opposite label. Third, the physician gave a final diagnosis, recorded as FHD, with a further confidence rating.
The pro-hoc design. The retrieved cases came with their verified ground-truth diagnoses attached, but the system offered no categorical advice and no probability score, abstaining entirely from interpreting the case in hand. The two same-label cases function as support for the physician's own hypothesis; the opposite-label case functions as a counterfactual, posed by the authors as the objection what if you were wrong, this would be the most similar case with the opposite label. Retrieval used Cosine similarity, selected because a previous user study by the same group found it the similarity metric most correlated with human similarity ratings.
Full findings
The support did not significantly improve accuracy, and the authors' argument is that this is the wrong yardstick for a system that deliberately declines to give advice. Pre-support accuracy was 78.8% against 80.9% post-support (two-proportion test p = .53, Z = -0.62, effect size .05), which is not significant. Sensitivity barely moved (.903 to .910, p = .84, effect size .02) and specificity rose more but still not significantly (.674 to .708, p = .52, effect size .08).
The substantive result is in the direction of the changes rather than their number. Only 16 of 288 decisions changed at all. Of those, 11 went from wrong to right and 5 from right to wrong, so errors fell from 61 to 55, a 10% reduction. The Number of Decisions Needed was 50 overall, meaning roughly fifty supported diagnoses would avert one error that would otherwise have been made. Technology Impact was slightly positive and, unlike the accuracy comparison, significantly so: the 95% confidence interval for the odds ratio excludes the line of no impact. The authors note that small effect sizes are normal when the effect of explanation is isolated from the effect of the AI's own advice, citing a comparable study where visual attribution maps attached to an 80% accurate system produced an effect size of 0.08.
Expertise reverses the expected pattern twice over. Unaided, residents outperformed specialists (.83, SD .06 against .76, SD .08), a large effect size of 1.03 though not significant (p = .5, T = 2.1); the authors attribute this to residents treating the exercise as a test of their skills and taking about 15%, or four minutes, longer. After support the gap almost closed, because the benefit went to the specialists: their effect size was small-to-moderate at .3 with an NDN of 6, and 60% of them improved, whereas no resident improved at all. The authors suggest fixation on the part of the residents.
Automation bias, measured as changes from a correct to an incorrect diagnosis, was two and a half times larger in residents than in specialists (0.028 against 0.011). Two-thirds of all decision changes concerned cases the physician had rated as complex, and two-thirds were for the better.