← Back to the framework

Experimental evidence of effective human-AI collaboration in medical decision-making

N = 21 Medical imagingClinical assessment / diagnosis

This study in the framework

Human Inherited
Expertise level Manipulated
Experts - juniorExperts - senior
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes High
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance Not reported
AI design Manipulated
Protocol Manipulated (AI-first, No AI (control))
AI stance Directive
Interactivity Static
XAI Not present
Number of AI advisors One
Research questions
  • RQ1: Are endoscopists influenced by the AI's opinion?
  • 2. Does the AI's advice lead to an improvement in diagnostic accuracy?
  • 3. Are endoscopists capable of selectively following the AI when it is correct, and conversely rejecting its opinion when it is incorrect (calibrating reliance)?
  • 4. How do individual features, specifically the endoscopist's level of expertise, modulate the AI's influence (e.g., does accuracy increase more for non-experts)?
In the authors' words

"Endoscopists were influenced by ai, but not erratically: they followed the ai advice more when it was correct than incorrect"

abstract

"This Bayesian-like rational behavior allowed the human-ai hybrid team to outperform both agents taken alone"

abstract

"experts were less able than non-experts to discriminate between good and bad ai advice"

results, effects of expertise
Experimental design

A multicentric within-subjects experiment with a mixed 2 by 2 design: treatment (no AI against AI) within subjects, expertise (experts against non-experts) between subjects. Twenty-one endoscopists each judged the same 504 lesions twice, giving 21,168 diagnostic observations.

Participants. Twenty-one endoscopists from Austria, Israel, Japan, Portugal and Spain. Ten were classed as experts, defined as at least five years of colonoscopy experience plus experience in optical biopsy with virtual chromoendoscopy; eleven as non-experts, defined as fewer than 500 colonoscopies performed. Task and instructions were in English. Approved by Comitato Etico Lazio 1 and reported following STROBE.

Stimuli. 504 short video clips, each showing one colorectal lesion, prospectively acquired in real colonoscopies during the CHANGE clinical study and captured in full length with no AI overlay. Ground truth was the histopathological diagnosis of each lesion.

Sessions. In session 1 the endoscopists diagnosed the lesions with the device running in detection mode only: GI Genius v3.0 in CADe modality dynamically drew a green box around lesions it detected, but offered no diagnosis. In session 2 they saw the same videos with the device in CADe plus CADx modality, adding a dynamically updating optical diagnosis that could change between frames across four displayed states: adenoma, non-adenoma, no-prediction, or analyzing.

Tasks per lesion. Categorise the lesion into five forced-choice options (adenoma, hyperplastic, SSL, carcinoma, uncertain), later collapsed to adenoma, non-adenoma or uncertain; then rate confidence on a four-point scale (very high, high, low, very low). Time from the start of the video to the first decision was recorded. No feedback was given at any point. In session 2 participants additionally reported their perception of the AI's overall opinion on the lesion (three options) and of the AI's confidence (four options).

Full findings

Endoscopists integrated the AI's opinion in proportion to its reliability, and the resulting hybrid team beat both agents alone. This is one of the clearest demonstrations of complementarity in the base, and it was achieved with a device that offered no explanation at all.

All four pre-registered predictions were supported. The AI exerted substantial influence (omega I = 3.05): for every three lesions on which endoscopist and AI disagreed in session 1, only one disagreement survived into session 2. Accuracy improved correspondingly, with the authors summarising it as seven lesions correctly evaluated with AI for every five correct without it.

The influence was discriminating rather than uniform. Endoscopists followed correct AI advice substantially more (effectiveness omega E = 3.48) than incorrect advice (1 divided by safety = 1.85). They did not escape the cost of bad advice entirely, since safety was below 1 and wrong AI opinions did drag some correct human diagnoses off target, but the gain from accepting good advice more than covered that loss. The net result is the complementarity claim.

The mechanism is Bayesian-like weighting of two reliabilities, and the authors demonstrate each component separately. Endoscopist confidence was strongly predictive of their own accuracy in both sessions. Their reading of the AI's dynamic output matched an algorithmic reading 81% of the time, rising to 94% once uncertain judgments are excluded, and their agreement with each other was similarly high (77%, rising to 94%). Their estimates of the AI's reliability predicted the AI's actual accuracy. Crucially, they then used both estimates: when their own confidence was high or the perceived AI confidence low, they held their position under disagreement; when their own confidence was low or the AI's perceived confidence high, they moved. Reliance was conditioned case by case on the two reliabilities rather than set as a global disposition.

Expertise moderated everything, and not in the direction one might assume. Non-experts were more influenced by the AI and gained more accuracy from it, which the authors attribute to experts having less headroom since their unaided accuracy already approached the AI's. But experts were also worse at telling good advice from bad: both effectiveness and safety were higher in non-experts. The authors read the experts' lower safety as a deliberate clinical asymmetry rather than a failure, since experts accepted incorrect adenoma advice more readily than non-experts, trading false positives for avoided false negatives, and in session 2 experts raised only their sensitivity while non-experts raised both sensitivity and specificity.

Confidence does not explain the expertise difference. Experts reported lower average confidence than non-experts in both their own judgments and the AI's, but the relative standing of the two was the same in both groups, with each rating the AI slightly above themselves. Confidence was predictive of real accuracy in both groups.

The AI's own accuracy depends on how its dynamic output is read, and the paper reports four figures rather than one: 84.9% using the endoscopists' interpretation and excluding uncertain outputs, 72.9% counting uncertain outputs as errors, 84.5% under algorithmic interpretation excluding uncertain, and 79.3% under algorithmic interpretation counting uncertain as errors.