← Back to the framework

More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice

N = 348 Income predictionRecidivism risk assessment

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes High
Stress level / time constraint Not reported
Task uncertainty Manipulated
Low (diagnostic)High (prognostic)
AI Ecosystem Designable
AI performance 70%-80%
AI design
Protocol Human-first (update)
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Feature importance · Narrative
Quality Correct
Number of AI advisors Manipulated
Count Manipulated (One, Three, More than three)
Agreement Manipulated (Unanimous, Divided)
Research questions
  • RQ1: How does the AI panel size shape human decision-making?
  • RQ2: How does the within-panel consensus of AIs affect human decision-making?
  • RQ3: How does the human-likeness of AI panels influence human decision-making?
In the authors' words

"High consensus fostered overreliance; a single dissent reduced pressure to conform; wide disagreement created confusion and undermined appropriate reliance."

abstract

"Conformity is a major reason for degraded integration: multiple advisors create pressure to follow the majority even when it is wrong."

introduction

"High consensus in human groups is known to strengthen conformity pressure."

discussion

"These comments suggest that increasing the number of AIs does not necessarily strengthen psychological reliance on AI."

results
Experimental design

Two studies using the judge-advisor system paradigm, the first between-subjects on panel size, the second between-subjects on advisor presentation.

Study 1, N = 260, split across three tasks and three panel sizes: income prediction with 30, 30 and 28 participants at panel sizes of one, three and five AIs; recidivism prediction with 28, 29 and 27; dating prediction with 26, 30 and 32.

Study 2, N = 88, held the panel at three AIs and varied whether the advisors were presented with a human-like or a robot appearance.

Tasks. Three binary prediction tasks, each participant assigned to one: income prediction from 7 features of the UCI Adult dataset, recidivism prediction from 8 features of COMPAS, and dating prediction from 11 features of a speed-dating dataset. Each participant completed 3 training trials and 50 test cases.

Procedure per trial. The profile was shown, the participant made an initial prediction and rated confidence on a 7-point scale, the AI advice was then displayed together with a natural-language explanation, an attention check required selecting the highlighted features, and the participant then gave a final prediction and rated confidence again. In multi-AI conditions the advisors' advice was presented sequentially, each with its own attention check.

AI systems. Decision trees sampled from random forests, each calibrated to 70% accuracy. Trees were drawn from Rashomon sets of 30 to 36 trees all achieving exactly 70% accuracy, so that panel members could disagree with one another while remaining equally accurate. This is the design device that makes within-panel consensus an independent variable rather than a confound with advisor quality.

Framing. Recidivism was explicitly framed as high-stakes: as the authors put it, recidivism prediction is framed as a high-stakes judicial task, which is widely recognized as contributing to underreliance on AI advice, and participants were reminded of its real-world significance. The other two tasks carried no such framing.

Measures. Decision accuracy as proportion correct; reliance through agreement fraction, switch fraction, RAIR, RSR and accuracy-wid; confidence change as the difference between the two 7-point ratings; and a 10-item questionnaire covering conformity pressure and perceived usefulness. Open-ended responses were collected.

Full findings

Three AI advisors beat one, five added nothing, and the reason more advisors stop helping is that agreement among them becomes conformity pressure rather than evidence.

On accuracy, the panel of three outperformed the single advisor on income prediction (M = .737 against .706, p = .012) and on dating prediction (Mdn = .68 against .64, p = .002). The panel of five was not better than three on either, reaching only marginal significance on dating (p = .064), and recidivism showed no panel-size effect at all (F(2,81) = 2.33, p = .104). The curve flattens after three.

The mechanism sits in RQ2 rather than RQ1. Within-panel consensus, not panel size, drives how advice is used: high consensus fostered overreliance, a single dissenting advisor reduced the pressure to conform, and wide disagreement produced confusion and information overload that undermined appropriate reliance. The authors are explicit that this imports a known human-group dynamic, high consensus in human groups is known to strengthen conformity pressure, into a setting where the advisors are machines and their agreement carries no independent evidential weight, since all panel members were equally accurate by construction.

Reliance did not scale with panel size. Neither agreement fraction nor switch fraction rose consistently with more advisors, and the RAIR and RSR decomposition was task-dependent rather than size-dependent: stronger reliance on AI in income prediction, stronger self-reliance in the recidivism and dating tasks. Confidence rose significantly after advice in almost every condition regardless of how many advisors gave it, which the authors read as the gain coming from receiving advice at all rather than from its quantity.

The qualitative data supports a non-monotonic account of autonomy. Participants with one advisor described consulting it and deciding themselves; participants with three described changing their answer when all three disagreed with them; the authors describe the relationship as shifting from autonomy to reliance and back to autonomy as panel size grows. In the recidivism and dating tasks participants stressed that the final decision was their own and the AI a reference only, which the authors attribute to ethical concern in the judicial domain and to algorithm aversion in the affective one.

Study 2 found that a human-like presentation of the three advisors raised perceived usefulness and perceived agency in some tasks without raising average conformity pressure, though individual differences were visible. The reporting for RQ3 is thinner than for RQ1 and RQ2, and the preprint does not give consolidated test statistics for either the consensus or the human-likeness analyses.