← Back to the framework

"When Two Wrongs Don't Make a Right" - Examining Confirmation Bias and the Role of Time Pressure During Human-AI Collaboration in Computational Pathology

N = 28 Medical imaging

This study in the framework

Human Inherited
Expertise level Experts - senior
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes High
Stress level / time constraint Manipulated
LowHigh
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance Not reported
AI design
Protocol AI-first · No AI (control)
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Example-based · Visual / saliency
Quality Correct
Number of AI advisors One
Research questions

H1a: Experts are more likely to accept inaccurate AI recommendations if their own prior assessment deviates similarly, demonstrating confirmation bias.

H1b: Confirmation bias manifests as pathologists weighting AI advice congruent with initial flawed judgments more heavily than incongruent system output.

H1c: Experts report greater confidence in decisions when AI advice aligns with initial erroneous beliefs versus incongruent advice.

H2: Time constraints increase the magnitude of confirmation bias.

In the authors' words

"This cognitive bias poses a significant challenge to the effectiveness of human-AI collaboration, as AI predictions cease to be independent from the medical expert due to a greater likelihood of accepting congruent recommendations."

discussion

"Both human and AI-generated errors may go unnoticed when coinciding error patterns lead to false confirmation, potentially impacting the quality of patient care."

discussion

"Conversely, time pressure appeared to weaken this relationship... while time pressure increases overall reliance on AI advice, it does not seem to exacerbate confirmation bias."

results, H2
Experimental design

Two-by-two factorial within-subjects design crossing AI assistance with time pressure, run in two phases separated by a two-week wash-out.

Participants. 31 pathology experts completed the first round and 28 continued to the second, which is the analysed sample: 25 pathologists, 2 residents and 1 non-physician staff member. Recruitment was through a professional network on a voluntary basis. A majority reported 15 or more years of experience. Exclusion was for failing to complete all slides in every condition, or for multiple entries by the same individual.

Task. Estimating tumour cell percentage on haematoxylin-and-eosin-stained tissue patches through a web interface with zoom. Each participant assessed 20 image patches taken from three open-access datasets, BreCaHad, Frei et al. and BreastPathQ.

Structure. Phase 1 was independent assessment with no AI. Phase 2, two weeks later, repeated the same patches with AI assistance. Within each phase the patches were split into two sets of ten, and half the slides in each segment were placed under time stress while the other half were not. Presentation order was randomised with a balanced 10x10 Latin square.

Time pressure manipulation. A 10-second countdown shown as a reverse progress bar shading from mint to blush, with a lock overlay blurring the image at 7.5 seconds. Participants could still submit after expiry but were discouraged from doing so. The no-pressure condition had no timer.

AI system. An FCOS-based object detection model trained on BreCaHad, producing a tumour cell percentage prediction together with prototype-based explanation showing exemplar tumour and non-neoplastic cells, and colour-coded cell detections overlaid on the image, rust red for neoplastic and teal for non-neoplastic. Participants could toggle the detection overlay. Model behaviour was deliberately realistic rather than uniformly good: on roughly half the patches detections were highly accurate, and on the remainder there were substantial false-positive and false-negative rates. Participants were not told any of this.

Measures. Confirmation bias was tested in two steps. Step 1 fitted linear mixed-effects models relating the distance between the AI prediction and the participant's own phase 1 baseline to the distance between their final estimate and the AI prediction. Step 2 ran a weighted-averaging simulation comparing the coefficients on the baseline estimate and on the AI prediction separately for congruent and incongruent advice. Alignment with advice was measured with a judge-advisor system value, JAS = |Est_AI − Est_B| / (|Est_AI − Pred_AI| + |Est_AI − Est_B|). Confidence was rated on a 5-point scale from not at all confident to completely confident.

Analysis. Linear mixed-effects models fitted with lme4 in R 4.4.0, p-values from lmerTest using Satterthwaite's approximation, paired one-tailed t-tests for the time-pressure contrasts, and Shapiro-Wilk tests for normality.

Full findings

Pathologists gave more weight to AI predictions that agreed with their own earlier judgement than to ones that disagreed, even when both were wrong, and time pressure did not make that worse.

The first-step model found a significant positive relationship between how far the AI prediction sat from the participant's own baseline and how far their final estimate ended up from the AI (beta = 0.61, p < .001), with an intercept of −0.72 meaning that when baseline and AI agreed the final estimate landed essentially on the AI prediction. The second step localises the asymmetry. Where AI advice was congruent with the baseline, the final estimate loaded 0.60 on the AI prediction and 0.25 on the baseline; where it was incongruent, the loading flipped towards the participant's own view, 0.43 on the AI and 0.47 on the baseline. The same pattern appears in the reliance and confidence measures: congruent predictions drew mean JAS 0.55 and mean confidence 3.87, incongruent ones 0.49 and 3.24. H1a, H1b and H1c are supported.

The clinical consequence is what makes this more than a bias demonstration. Because the model erred substantially on about half the patches and the experts erred too, agreement between them is not independent corroboration. An AI prediction that happens to share the pathologist's error is precisely the one they will weight most and feel most confident about, so the two error sources reinforce rather than cancel.

H2 was not supported. Time pressure slightly weakened the confirmation-bias slope (beta = 0.65 without pressure against 0.56 with it, both p < .001) while raising overall reliance on the AI (JAS M = 0.49, SD = 0.13 without pressure against M = 0.55, SD = 0.12 with it; paired t(27) = −2.80, p = .005). Under a 10-second clock experts leaned on the AI more but discriminated less between advice that agreed with them and advice that did not. The authors read this as a shift from selective information processing to heuristic acceptance: with no time to compare the recommendation against one's own reading, there is nothing for congruence to act on. Time pressure and confirmation bias are therefore two distinct failure modes rather than one amplified, and an intervention against one may not touch the other.