Does More Advice Help? The Effects of Second Opinions in AI-Assisted Decision Making
N = 1280 Sentiment analysisThis study in the framework
Held constant. A RoBERTa model fine-tuned on 5,000 IMDB reviews reached 0.776 accuracy on a 500-review held-out test set and 0.750 on the 20 reviews used in the experiments, correct on 15 and incorrect on 5. The same model served every treatment in all three experiments. The authors state they deliberately avoided training a high-accuracy model so that over-reliance and under-reliance could both be observed with adequate data. No per-case confidence was displayed; subjects saw a bare binary label.
- RQ1: How do second opinions from human peers affect decision-makers' reliance on the AI model (e.g., over-reliance, under-reliance, and appropriate reliance) and influence their decision-making performance (e.g., decision accuracy, time and confidence)?
- RQ2: Do the impacts of second opinions on decision-makers' reliance on AI and performance in AI-assisted decision making change, when the second opinions are claimed to be solicited from another AI model rather than human peers?
- RQ3: How does having the option to solicit second opinions affect decision-makers' reliance on the AI model and influence their performance in AI-assisted decision making?
"the presence of second opinions significantly reduces people's reliance on the AI no matter whether the second opinions align with the AI model's recommendation or not"
results, experiment 1 exploratory analysis
"the source of the second opinions does not appear to play a major role in influencing how decision-makers would utilize the second opinions"
results, experiment 2
"the mere presence of second opinions from peers is not sufficient for encouraging people to engage in deep deliberative thinking on each decision-making task"
discussion
Experimental design
Three pre-registered, randomised, between-groups experiments on Amazon Mechanical Turk, restricted to United States workers, one participation per worker, and with workers from earlier experiments barred from later ones. All three used the same task and the same AI model.
Task. Each subject judged the same 20 IMDB movie reviews, each controlled to between 280 and 300 words, as positive or negative in sentiment. Ground truth came from the reviewer's own star rating, with ambiguous mid-range ratings excluded from the source dataset. The AI's binary prediction was displayed alongside the review before the subject decided; no independent human judgement was recorded first, and neither an explanation nor a confidence score was shown. Subjects rated their confidence on each task on a 7-point scale. The exit survey collected demographics and asked subjects to estimate the AI's, the peers' and their own accuracy across the 20 tasks. Base payment was $1.20, with a 5-cent bonus per correct decision paid only if the subject's accuracy exceeded 65%, for a maximum bonus of $1. An attention check gated inclusion and a qualification task gated entry.
Second opinions. Peer judgements were collected in advance from a pilot of 34 crowd workers who labelled the same 20 reviews independently, without seeing the AI. Each worker's level of agreement with the AI was computed, and three sets of three workers were formed: high agreement, medium agreement and low agreement, averaging 76.67%, 50% and 30% agreement respectively. On each task a worker was drawn at random from the assigned set and that worker's judgement was replayed to the subject as the second opinion.
Experiment 1 (N = 428) had four treatments: control with no second opinion, and high, medium and low agreement second opinions presented on every task.
Experiment 2 (N = 516) had five treatments: a control with no second opinion, plus a 2 by 2 factorial crossing the agreement level of the second opinion (high or low) with its stated source (another crowd worker or another AI model trained with a different algorithm). Content was held identical within an agreement level so that only the stated source varied.
Experiment 3 (N = 336) kept the control, high agreement and low agreement treatments but presented the second opinion only when the subject clicked a Request button, rather than automatically. A three-item cognitive reflection test was added to the exit survey.
Outcomes were operationalised as overall reliance (probability the subject's decision matches the AI's prediction), over-reliance (matching the AI when the AI is wrong), under-reliance (departing from the AI when the AI is right), appropriate reliance (equivalent here to decision accuracy), decision time, and confidence in correct and in incorrect decisions computed separately at task level. Analysis was one-way ANOVA with Tukey HSD post-hoc for reliance and time, repeated-measures ANOVA with Bonferroni-corrected paired t-tests for confidence, two-way ANOVA for the factorial in experiment 2, and mixed-effect logistic regressions with subject and task as random effects for the task-level exploratory analyses. Experiment 3's subgroup analysis used propensity score matching on demographics, prior programming knowledge and cognitive reflection score to pair soliciting subjects with comparable controls.
Full findings
Presenting a second opinion on every task buys a reduction in over-reliance at the price of an increase in under-reliance, and the two cancel: appropriate reliance, which here is decision accuracy, does not improve. In experiment 1 overall reliance fell significantly in all three second-opinion treatments relative to control (F(3,8556) = 28.31, p < 0.001; Cohen's d 0.17 to 0.28, rising as peers disagreed more). Over-reliance fell significantly in every treatment (F(3,2136) = 17.27, p < 0.001; d 0.26 to 0.44) while under-reliance rose significantly in every treatment (F(3,6446) = 13.84, p < 0.001; d 0.14 to 0.22). Appropriate reliance showed no significant difference. Decision time rose significantly (F(3,8122) = 12.06, p < 0.001) and confidence in correct decisions rose in the low agreement treatment relative to both control (p = 0.003, d = 1.08) and high agreement (p = 0.002, d = 1.64); confidence in incorrect decisions showed no significant pairwise differences.
The mechanism the authors identify is that disagreement drives the whole effect and does so indiscriminately. The authors' preferred reading is that subjects use peer-AI agreement as an heuristic for the AI's overall accuracy rather than engaging analytically with the individual case.
The stated source of the second opinion did not matter. In experiment 2 the two-way ANOVA found significant main effects of agreement level on overall reliance (p < 0.001), over-reliance (p = 0.002) and under-reliance (p = 0.038), but no significant main effect of source and no significant interaction on any reliance measure. Second opinions attributed to another AI model behaved like second opinions attributed to human peers. Experiment 2 also produced the study's one accuracy cost: subjects in the high agreement-AI source and the low agreement-human source treatments showed significantly lower appropriate reliance than control (p < 0.05).
Making solicitation optional does not by itself fix the problem. Across all subjects in experiment 3, overall reliance fell (F(2,6717) = 14.25, p < 0.001) and under-reliance rose (F(2,5037) = 11.91, p < 0.001) relative to control, while differences in over-reliance and appropriate reliance were not significant. The favourable result is confined to an exploratory subgroup. Only 34.0% of subjects in the high agreement treatment (31 subjects) and 38.8% in the low agreement treatment (44 subjects) ever requested a second opinion. Among the 31 high agreement soliciters compared with propensity-matched controls, over-reliance fell significantly (0.61 to 0.47, p = 0.009) with no significant increase in under-reliance, and appropriate reliance rose from 0.69 to 0.73, a difference that was not statistically significant. Among the 44 low agreement soliciters, over-reliance fell (p = 0.003) but under-reliance rose (p = 0.003) and appropriate reliance did not change.