← Back to the framework

Impact of a deep learning assistant on the histopathologic classification of liver cancer

N = 11 Clinical assessment / diagnosisMedical imaging

This study in the framework

Human Inherited
Expertise level Manipulated
Experts - seniorExperts - junior
Team composition Individual
Task Inherited
Difficulty Manipulated
LowMediumHigh
Stakes High
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 80%-90%
AI design
Protocol Human-first (update) · No AI (control) · Confidence display
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Visual / saliency
Quality Correct
Number of AI advisors One
Research questions
  • RQ1. Does assistance change the diagnostic accuracy of pathologists distinguishing hepatocellular carcinoma from cholangiocarcinoma?
  • RQ2. Do pathologist experience level and case difficulty affect diagnostic accuracy, with and without assistance?
  • RQ3. How does the correctness of the model's prediction affect the pathologist's final diagnosis?
In the authors' words

"In the assisted state, model accuracy significantly impacted the diagnostic decisions of all 11 pathologists. As expected, when the model's prediction was correct, assistance significantly improved accuracy (p = 0.000, OR = 4.289), whereas when the model's prediction was incorrect, assistance significantly decreased accuracy (p = 0.000, OR = 0.253), with both effects holding across all pathologist experience levels and case difficulty levels."

abstract

"These results suggest that the pathologists might have relied more heavily on the model's output for difficult cases."

discussion

"The finding that inaccurate model predictions can have a strong negative impact, even on subspecialty pathologists with particular expertise at the diagnostic task in question, raises concerns about the unintended effects of decision support tools, such as automation bias."

discussion

"The observation that the combination of the model and pathologist outperformed both the model alone, and the pathologist alone, suggests that the model and the pathologist are complementary, rather than parallel, with respect to the features each uses to arrive at a correct diagnosis, and that the model should be used to augment, rather than replace, the pathologist."

discussion
Experimental design

Prospective within-subjects crossover study. Eleven pathologists diagnosed the same independent test set of 80 whole-slide images of hematoxylin and eosin stained primary liver tumour resections, 40 hepatocellular carcinoma and 40 cholangiocarcinoma, drawn from the Stanford University Medical Center archive.

Each pathologist read all 80 slides twice, in two sessions separated by a washout period of at least two weeks, following College of American Pathologists guidance on short-term memory bias. In each session half the slides were read with the assistant and half without; at the second reading the assistance status of every slide was reversed. Pathologists were randomised to begin either with or without assistance. Slides were presented in the same sequence in both readings, in eight blocks of ten preceded by a four-slide practice block. Pathologists were blinded to original diagnoses, clinical histories and follow-up.

On unassisted cases the pathologist read the slide in a whole-slide viewer and recorded a diagnosis. On assisted cases the pathologist first selected one or more tumour regions of interest, saved them as image patches at times ten objective magnification, uploaded them to a browser-based tool, and received the model's predicted probability for each diagnosis together with class activation maps highlighting the image regions most consistent with each. The final diagnosis was recorded after viewing this output. Assistance was therefore semi-automated: the model ran only on patches the pathologist chose.

The pathologists were classified into four experience subgroups: gastrointestinal subspecialists (n = 3), non-GI subspecialists (n = 3), trainees (n = 3), and pathologists not otherwise classified (n = 2). Case difficulty was operationalised through tumour grade, from grade 1 well-differentiated to grade 3 poorly differentiated. Analysis used mixed-effect logistic regression with fixed effects for experience level and case difficulty and random effects for pathologist and slide.

Full findings

Assistance did not significantly change mean accuracy across all eleven pathologists: 0.898 unassisted against 0.914 assisted (p = 0.184, OR = 1.281). Restricting to the nine pathologists of well-defined experience level, assistance did significantly improve accuracy (p = 0.045, OR = 1.499). Both experience level and case difficulty independently predicted accuracy: non-GI subspecialists (OR = 0.204) and trainees (OR = 0.299) were less likely than GI subspecialists to be correct, and grade 3 tumours were harder than grade 1 (OR = 0.157).

The central finding is the conditional effect of model correctness. When the model was correct, assisted pathologists had 4.3 times the odds of a correct final diagnosis compared with the same pathologists unassisted (OR = 4.289). When the model was wrong, they had less than one third the odds (OR = 0.253). Both effects held across every experience level and every case difficulty level, including among GI subspecialists reading within their own subspecialty.

The damage was concentrated in hard cases, which is where reliance was heaviest. On grade 3 cases, unassisted accuracy was 0.76; with a correct model prediction assisted accuracy rose to 0.947, and with an incorrect one it fell to 0.310. On grades 1 and 2, unassisted accuracy was 0.917, rising to 0.982 with a correct prediction and falling to 0.654 with an incorrect one.