← Back to the framework

Rams, hounds and white boxes: human-AI collaboration protocols in medical diagnosis

N = 56 Medical imagingClinical assessment / diagnosis

This study in the framework

Human Inherited
Expertise level Manipulated
Experts - juniorExperts - senior
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes High
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 80%-90% · 70%-80%
AI design Manipulated
Protocol Manipulated (AI-first, Human-first (update))
AI stance Directive
Interactivity Static
XAI Manipulated
Presence Manipulated
Type None · Visual / saliency · Narrative
Quality Manipulated
Number of AI advisors One
Research questions
  • RQ1. Do the main factors under analysis (i.e., human-first vs AI-first protocol, provisioning of explanations, readers' expertise) have any effect on accuracy?
  • RQ2. Do the individual HAI-CPs (i.e. the interactions between presentation order and provisioning or not of explanations) under analysis have any effect on accuracy?
  • RQ3. Did the provisioning of AI support and/or explanations allow the readers to improve their accuracy, in comparison with their unsupported accuracy, or to surpass the accuracy of the AI support?
  • RQ4. Were there any effects (in terms of increased risk of automation-related biases) of computer support on reliance patterns and behavioral correlates of trust?
In the authors' words

"we confirm the utility of AI support but find that XAI can be associated with a 'white-box paradox', producing a null or detrimental effect"

abstract

"the order of presentation matters: AI-first protocols are associated with higher diagnostic accuracy than human-first protocols, and with higher accuracy than both humans and AI alone"

abstract

"the availability of explanations increased the chance that the advice of the AI was trusted by the clinicians. Furthermore, this increase in trust due to the XAI support was observed irrespective of whether the AI advice was right or wrong"

discussion

"by recalling the metaphor presented in the title, using AI as a ram is better than using it as a hound"

discussion
Experimental design

Two user studies with practising clinicians, each a 2 by 2 factorial crossing presentation order (human-first, AI-first) with explanation availability (yes, no). The authors call each cell a human-AI collaboration protocol, HAI-CP, and name the two order classes rams (AI-first) and hounds (human-first). Both studies ran as online LimeSurvey questionnaires. The AI was simulated in both, which is what allowed accuracy and explanation content to be fixed by design.

Knee MRI study. 12 board-certified radiologists from Italian hospitals, 8 of higher expertise (subspecialists) and 4 of lower expertise (specialists), each reading 240 knee MRI exams drawn from the MRNet dataset and shown in random order with axial, sagittal and coronal views. For each case the reader judged whether a ligament abnormality, another knee abnormality, or neither was present. Both factors were within-subject: 120 of the 240 cases were AI-first and 120 human-first, and within each half 60 carried an explanation and 60 did not. Explanations were GradCAM activation maps highlighting the MRI regions the system treated as most relevant. Expertise was compared between-subjects. Four protocols result: AI-FHD, AI-XAI-FHD, HD1-AI-FHD, HD1-AI-XAI-FHD, where HD1 is the recorded first human decision and FHD the final one.

ECG study. 44 readers from the University Hospital of Siena, 25 cardiology residents (novices) and 19 specialists (experts), each annotating 20 ECG cases selected by a cardiologist from the ECG Wave-Maven repository on the basis of their complexity characteristics; 1352 responses in total. Here presentation order was between-subjects, 21 readers in human-first and 23 in AI-first, while explanation availability stayed within-subject. Human-first readers gave a free-text diagnosis, saw the AI diagnosis, could revise, then saw the textual explanation and gave a final diagnosis. AI-first readers saw the AI diagnosis alongside the trace, gave their own, then saw the explanation and could revise. To avoid negative priming the first five cases always carried a correct diagnosis and a correct explanation. Explanations were presented to participants as automatically generated but were in fact written by a cardiologist, and 40% were deliberately incorrect or not fully pertinent.

In both studies the first human decision in the human-first arm was retained as an unsupported baseline against ground truth (MRNet labels, ECG Wave-Maven gold standard).

Full findings

Showing the AI first beat making the clinician commit first, in both studies and against both baselines; explanations were null or harmful, and did their harm through trust rather than through accuracy.

Presentation order. In the MRI study AI-first protocols were significantly more accurate than human-first ones (p = .033, RBC .46, medium). In the ECG study the same comparison was significant with a large effect (p < .001, RBC .90). The ECG figures show the size of it: AI-first reached 83% for novices and 82% for experts, while human-first without explanations reached 63% and 68%, against unsupported baselines of 45% and 66%. AI-first protocols were significantly better than every other protocol and than the unsupported readers, with no distinction between novices and experts; human-first protocols were not significantly better than unsupported experts.

Complementarity. AI support beat not only the unaided readers but the AI alone, in the MRI study (p = .001) and the ECG study (p < .001), and this held despite AI accuracy of 80% and 70% respectively, at or below the average reader. The authors treat this as one of the few credible demonstrations of complementarity, and note it supports the claim that good hybrid performance is achievable with decision support weaker than the humans it assists.

Explanations, the white-box paradox. The main effect of explanation availability was not significant in either study, though the effect sizes were not small (MRI p = .095, RBC .62; ECG p = .081, RBC .55). The direction depended entirely on the protocol. Under AI-first, explanations were slightly positive or negligible (MRI 81% to 83% for lower expertise, 85% to 86% for higher; ECG flat at 82-83%). Under human-first in the MRI study they were clearly detrimental: accuracy fell from 84% to 73% for lower-expertise radiologists and from 85% to 80% for higher-expertise ones, taking the lower-expertise readers below their own unsupported baseline of 77%. Under human-first in the ECG study they were slightly positive (63% to 67% for novices, 68% to 69% for experts).

The mechanism is trust, not comprehension. The reliance-pattern analysis shows explanations moving clinicians towards accepting advice regardless of whether it was right. In the MRI study the XAI-supported human-first protocol had significantly lower relative beneficial distrust and relative beneficial over-distrust than the same protocol without explanations, meaning higher automation bias and higher automation complacency. The ECG study showed the same increase in automation bias, with lower conservatism bias. The authors state the increase in trust from XAI "was observed irrespective of whether the AI advice was right or wrong".

Expertise. Expertise separated the readers before support and stopped separating them after it. Unsupported ECG accuracy differed significantly between novices and experts (p = .009, RBC .74), but no supported protocol showed a significant expertise difference; the authors read this as AI support pulling less-expert readers up to expert level, attributing it to lower algorithmic aversion among the less experienced. In the MRI study no expertise comparison reached significance, though effect sizes were medium-to-large, and expertise was dropped from later analyses for lack of power.