← Back to the framework

Beyond AI Advice — Independent Aggregation Boosts Human-AI Accuracy

N = 1229 Sentiment analysisMedical imagingRecidivism risk assessmentDeception detectionImage classificationOther

This study in the framework

Human Inherited
Expertise level Mixed
Team composition Manipulated
Composition Manipulated (Individual, Group, Independent)
Size 2 humans plus AI, tiebreak only (synthesised by resampling)
Mode Independent
Task Inherited
Difficulty Not reported
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty Not reported
AI Ecosystem Designable
AI performance <60% · 60%-70% · 70%-80% · 80%-90%
AI design Manipulated
Protocol Manipulated (Human-first (update), No AI (control))
AI stance Directive
Interactivity Not reported
XAI Manipulated
Presence Manipulated
Type None · Feature importance · Example-based · Visual / saliency
Quality Correct
Number of AI advisors One
Research questions
  • RQ1. Does the hybrid confirmation tree, in which human and AI judge independently and a second human breaks ties, produce more accurate decisions than AI-as-advisor, in which the human sees the AI's recommendation and accepts or rejects it?
  • RQ2. Does that advantage survive when the AI-as-advisor arrangement is strengthened with explanations?
In the authors' words

"The HCT outperforms the AI-as-advisor approach because people cannot discriminate well enough between correct and incorrect AI advice."

abstract

"By placing the human decision after the AI's output, the onus is on the human to adequately distinguish between accurate and inaccurate AI advice. Yet humans struggle to do so."

introduction

"The tiebreaker in the HCT resolved conflict differently. When the AI judgment was correct, the tiebreaker agreed with it in 71% of cases—twice as frequently as the 34% uptake of accurate AI advice."

results

"Humans cannot discriminate well enough between accurate and inaccurate AI advice and are conservative advice-takers—that is, they rarely change initial judgments when disagreeing with AI systems."

discussion
Experimental design

Secondary re-analysis and simulation across published datasets. No new data was collected.

Material. Ten primary datasets containing over 41,000 judgments by 1,229 human decision-makers across 3,220 cases, drawn from medical diagnostics, misinformation detection, sentiment classification, deception detection and rearrest prediction. Per dataset, with participant counts: Bansal et al. 2021 beer reviews (92) and book reviews (91), sentiment classification, AI accuracy 84%; Chanda et al. 2024 skin cancer (109), 80% balanced accuracy; DeVerna et al. 2024 news headline accuracy (241), ChatGPT 3.5 at 52%; Fogliato et al. 2021 criminal rearrest (271), lasso logistic regression at 67%; Groh et al. 2022 deepfake detection (304), 78%; Lai and Tan 2019 deception detection in hotel reviews (80), support vector machine at 87%; Reverberi et al. 2022 colonoscopy lesions split into experts (10) and non-experts (11), 79%; Vodrahalli et al. 2022 skin cancer biopsy decisions by dermatologists (20), ResNet-18 at 83%.

A second set of four datasets carried 16 explainable-AI conditions, an additional 50,390 decisions by 1,423 humans across 516 cases, used to test whether explanations close the gap.

The two protocols compared. AI-as-advisor is the arrangement as run in each source study: the AI recommends and the human accepts or rejects. The hybrid confirmation tree, HCT, is constructed by the authors: "In the HCT, a human and an AI system make initial decisions independently. If they agree, the choice is accepted; if they disagree, a second human is consulted to break the tie." The second human is drawn by resampling from the same participant pool, so the HCT arm is synthesised rather than run.

Task requirements. Every source task was reduced to a binary judgement, with slider and confidence responses binarised where the original study used them, so that accuracy is comparable across domains.

Measures. Decision accuracy of the resulting choice under each protocol. Uptake of AI advice split by whether the advice was correct, which decomposes into acceptance of correct advice and rejection of incorrect advice. Frequency with which the HCT triggers a tiebreak, which is its operating cost.

Analysis. Bayesian generalised linear mixed models with cases nested in datasets, contrasts summarised by 95% highest density intervals, a region of practical equivalence of plus or minus one percentage point, and probability of direction and of practical significance. A signal detection model with equal variance was fitted to the human's discrimination of correct from incorrect advice, yielding d' for discrimination and c for response bias, and then used to predict which protocol should win in each comparison.

Moderator analyses. The HCT advantage was estimated separately for low, mid and high performers within the five datasets whose design permits a within-subject split, and separately for cases where the AI was correct and where it was incorrect.

Full findings

Taking the human's judgement before the AI's, and settling disagreements with a second human rather than asking the first to adjudicate the machine, beat the standard advisor arrangement in all ten datasets, and the reason is that people cannot tell good AI advice from bad.

The hybrid confirmation tree outperformed AI-as-advisor in every dataset, by between 0.2 and 6.6 percentage points, pooling to 4.45 points (95% HDI 3.73 to 5.27). Practical significance reached 100% in every dataset except beer reviews, where it was 22.9%. Explanations did not rescue the advisor arrangement: across 16 explainable-AI conditions in four datasets, the HCT was better in 11, comparable in 2 and worse in 3, and all three losses came from deception detection, where unaided human accuracy sits at 51.1%, essentially chance. That exception is the boundary condition of the result: independent aggregation only helps if the independent human signal carries information.

The mechanism is a failure of discrimination compounded by conservatism. In the advisor arrangement, when human and AI disagreed, participants adopted correct AI advice in only 34% of cases while rejecting incorrect advice in 80%. The same disagreements handed to a fresh second human resolved in favour of correct AI judgements 71% of the time, and against incorrect ones 47% of the time. The signal detection model reads this as low d', people cannot separate accurate from inaccurate advice, plus a conservative criterion c greater than zero, people default to their own initial judgement. The model predicted the winning protocol in 37 of 41 comparisons, 90.2%.

The advantage is not free and is not uniform. The HCT is by construction better when the AI is right and worse when the AI is wrong: against AI-as-advisor it gains 10.4 points (HDI 9.68 to 11.15) on correct-AI cases and loses 11.89 points (HDI −10.72 to −13.15) on incorrect-AI ones. It wins overall because correct AI cases are more common and because the first human's rejection of correct advice is the larger of the two errors in practice. It also costs a second opinion on 22% (book reviews) to 49% (deception detection) of cases.

The benefit falls off with the decision-maker's own ability, which is the finding most relevant to deployment: 8 percentage points for low performers (HDI 6.9 to 9.2), 3.2 for mid performers (HDI 1.9 to 4.6), and 1.9 for high performers (HDI 0.6 to 3.2, practical significance 92%). Independent aggregation is therefore most valuable exactly where advisor-style reliance is most likely to be uncritical, and its margin narrows to near-nothing for the strongest individual judges.