← Back to the framework

Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment

N = 326 Recidivism risk assessment

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Manipulated
Composition Manipulated (Individual, Group)
Size 3
Mode Interdependent (deliberating)
Task Inherited
Difficulty Not reported
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty High (prognostic)
AI Ecosystem Designable
AI performance <60%
AI design
Protocol AI-first
AI stance Directive
Interactivity Static
XAI Not present
Number of AI advisors One
Research questions

RQ: How groups' behaviour and performance in AI-assisted decision making compare with those of individuals.

In the authors' words

"compared to individuals, groups rely on AI models more regardless of their correctness, but they are more confident when they overturn incorrect AI recommendations"

abstract

"groups make fairer decisions than individuals according to the accuracy equality criterion, and groups are willing to give AI more credit when they make correct decisions"

abstract

"we conjecture that groups are more likely to accept the AI recommendation when all members have weak but disagreeing opinions, but they are more likely to confidently reject the AI recommendation if some member has a strong opinion that disagrees with AI"

discussion, section 5.1
Experimental design

A pre-registered, randomised, two-treatment between-groups experiment on Amazon Mechanical Turk, run in two phases on different days, comparing individuals against three-person groups on the same AI-assisted task.

Task. Recidivism risk assessment. Each trial presented a criminal defendant's profile with eight features covering demographics (gender, age, race), criminal history (counts of prior non-juvenile, juvenile misdemeanour and juvenile felony offences) and the current charge (charge issue, charge degree), together with the AI's binary prediction of whether the defendant would reoffend within two years. Profiles came from the COMPAS dataset of Broward County, Florida defendants from 2013 to 2014, in the version shared by Dressel and Farid. The domain was chosen for three stated reasons: collective decision making is already normal in criminal justice, AI assistance is already deployed there, and the COMPAS algorithm's documented racial disparity permits a fairness analysis.

The AI. The COMPAS algorithm itself, not a model trained for the study. Its recidivism risk score, on a scale of 1 to 10, was thresholded at 5 or above to produce the binary prediction participants saw. No confidence and no explanation were displayed.

Treatments. In Individual-AI, the participant reviewed the profile and the COMPAS prediction and made a binary prediction alone. In Group-AI, three participants saw the same profile and the same prediction and used an embedded chatroom to discuss the case until all members selected the same prediction, which they then confirmed as the group's final decision. A risk analysis bot posted the AI's recommendation as the first message in every group chat by default. Participants idle for one minute were prompted; those inactive for two minutes were removed from the group.

Phase 1, recruitment and screening. A recruitment HIT collected a demographic survey and nine unassisted recidivism predictions in randomised order with no feedback. One of the nine was an attention check; only those who passed were invited to Phase 2. The other eight profiles were balanced on defendant race, charge degree and true recidivism status while being nearly identical on all other features, which allowed each participant's cognitive style to be quantified as the difference in positive-prediction rates between Black and White defendants and between felony and misdemeanour defendants.

Phase 2, the experiment. Fifteen AI-assisted tasks: nine practice then six formal. The practice tasks gave immediate feedback on the true recidivism outcome, followed by a summary page reporting the accuracy of both COMPAS and the participant, so that participants could calibrate a strategy. One practice task was an attention check; the other eight were balanced across four Black and four White defendants, on which COMPAS scored 62.5%, close to its dataset-wide accuracy, with a higher false positive rate on Black defendants. Participants were then randomly assigned to a treatment. Group-AI participants waited in a lobby for two others and chose an animal avatar for anonymity; anyone waiting more than five minutes was redirected to Individual-AI. The six formal tasks carried no feedback. They were assembled per participant or group by drawing one profile from each of four prepared subsets crossing race with true recidivism status, plus one pair of twin profiles identical on every feature and outcome except race, so that the six tasks were balanced on race with equal reoffending base rates across races.

Measures. Decision accuracy as the fraction of the six formal tasks answered correctly. Overall reliance as the fraction of formal tasks where the final prediction matched the AI's; over-reliance as agreement restricted to tasks where the AI was wrong; under-reliance as disagreement restricted to tasks where the AI was right. Decision confidence self-reported on a 5-point scale after every formal task, separately by each group member for the group's decision. Understanding of the AI as the Pearson correlation between participants' rated feature importance and the true importance, the latter being the absolute correlation between each feature and the AI's predictions. Decision fairness through five metrics comparing Black against White defendants: positive prediction difference, twin case prediction difference, accuracy difference, false positive rate difference and false negative rate difference, the first three treated as primary. Interaction fairness through two further metrics, the difference in reliance on positive and on negative AI predictions across defendant race. Accountability through a pie chart in the exit survey allocating credit for correct decisions and responsibility for incorrect ones between the participant, the AI and, for groups, teammates, normalised by the equal share so that values above 1 indicate more than an equal share.

Full findings

Groups relied on the AI more than individuals did, and they did so whether the AI was right or wrong, so the extra reliance bought them nothing in accuracy. Overall reliance rose from 0.60 (SD 0.26) for individuals to 0.73 (SD 0.24) for groups, t(184) = -3.47, p < 0.001. Decomposed, under-reliance fell from 0.37 (SD 0.34) to 0.24 (SD 0.30), t(184) = 2.86, adjusted p = 0.009, which is desirable; but over-reliance rose from 0.53 (SD 0.36) to 0.68 (SD 0.33), t(180) = -2.89, adjusted p = 0.010, which is not. Deliberation shifted reliance without calibrating it.

Decision accuracy did not differ: 57.8% for groups against 55.3% for individuals, t(184) = -0.75, p = 0.452. Understanding of the AI did not differ either, and was close to nil in both arms, with the correlation between perceived and true feature importance averaging 0.05 for individuals and 0.10 for groups, p = 0.344. Two of the six aspects therefore returned nulls, and the understanding null is doubly informative: neither individuals nor groups built a usable model of what drove COMPAS's predictions.

Confidence showed one targeted difference. Aggregate confidence did not differ by treatment for either correct or incorrect decisions. But conditioning on both correctness and agreement with the AI revealed that individuals became much less confident whenever they declined to follow the AI, whereas groups did not. In the specific cell where participants overturned an incorrect AI recommendation and were right to do so, groups were more confident than individuals, 4.09 (SD 0.85) against 3.81 (SD 0.88), adjusted p = 0.047, with no difference in the other three cells. Groups are therefore not blindly deferential; they follow the AI more often overall but overturn it with more conviction when they do overturn it.

Accountability shifted toward the machine when people worked together, and only in success. For correct final decisions, group members gave themselves less credit than individuals did, 1.01 (SD 0.19) against 1.10 (SD 0.29), t(1115) = 5.31, adjusted p < 0.001, and gave the AI more, 0.96 (SD 0.30) against 0.89 (SD 0.30), t(1115) = -3.30, adjusted p = 0.002. For incorrect decisions there was no difference in what either party was assigned. Disaggregating further, the effect came entirely from trials where the AI's recommendation was correct. Groups share the credit with the AI when things go well and do not shift extra blame to it when they do not.

The chat logs supply the mechanism. Groups routinely opened deliberation by stating the AI's recommendation, used it to justify or refute a position, invoked it to back up an opinion, and, revealingly, fell back on it as a tiebreaker when members could not persuade one another. Some members re-posted the recommendation mid-discussion, effectively appointing themselves defenders of the AI's position; the authors suggest others may then feel pressure to agree in order to reach the required consensus. The requirement to converge, not the AI itself, is what converts disagreement into deference.