Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's Advocate
N = 350 Recidivism risk assessmentThis study in the framework
Held constant, and deliberately biased. RiskComp is a random forest classifier trained on a COMPAS sample from which Black defendants with low prior-crime counts were filtered out. Overall accuracy is 66%, but that figure conceals the structure the study depends on: 62.5% on in-distribution cases and 48%, essentially chance, on the out-of-distribution cases, meaning Black defendants with few prior crimes.
- RQ1: Can LLM-powered devil's advocate help groups utilize AI assistance more appropriately?
- RQ2: How do the target of objection and interactivity of the LLM-powered devil's advocate affect the appropriateness of groups' utilization of AI assistance?
- RQ3: How does LLM-powered devil's advocate affect groups' utilization of AI assistance in the in-distribution and out-of-distribution decision making cases, respectively?
- RQ4: How does LLM-powered devil's advocate affect groups' perceptions of the group processes?
"The devil's advocate is often asked to argue for a position that is different from the accepted norm within the group (e.g., the majority opinion among group members), for the sake of provoking debate, testing the strength of the opposing arguments, and forcing the group to explore more diverse perspectives."
background
"Despite participants in this treatment having the highest decision accuracy among participants in all treatments, they reported the lowest level of self-perceptions of decision making performance... the constant argumentation and debate brought up by the devil's advocate lead to some level of discomfort or discord within the group, even if it may in fact improve group performance."
discussion
"The exceptional conversational capabilities exhibited by the state-of-the-art large language models (LLMs) appear to offer a viable solution to fully release the potential of the devil's advocacy technique in enabling groups' appropriate utilization of AI assistance."
introduction
Experimental design
Between-groups design with five treatments, crossing the target of the devil's advocate's objection with its interactivity, plus a no-advocate control.
Task. Recidivism risk prediction on COMPAS defendant profiles described by 8 attributes covering demographics and criminal history. Each participant made an independent prediction, then saw the AI recommendation, then discussed with the group in a chat, then finalised their own prediction; the group decision was the majority. Eight defendant profiles were used in the formal tasks, two from each combination of race and prior-crime count.
Conditions. Control, with no devil's advocate. Static-AI, a non-interactive advocate challenging the AI recommendation. Static-Majority, a non-interactive advocate challenging the majority opinion in the group. Dynamic-AI, an interactive advocate challenging the AI recommendation. Dynamic-Majority, an interactive advocate challenging the majority opinion.
AI system. RiskComp, a random forest classifier trained on a deliberately biased sample of COMPAS from which Black defendants with low prior-crime counts had been filtered out. Overall accuracy 66%, 62.5% on in-distribution cases and 48% on the out-of-distribution cases, which are Black defendants with few prior crimes. The bias is the point: it creates a subpopulation on which the model is at chance while looking competent overall.
Devil's advocate. Built on GPT-3.5-turbo at temperature 1. The non-interactive version generated three critiques of under 20 words each, presented at the start of discussion. The interactive version used a three-step pipeline, classifying the intent of a participant's message, then its stance, then generating a critique in response, so that it argued back through the discussion.
Incentives. USD 0.20 base for phase 1 and USD 2.40 for phase 2, a USD 0.25 per minute lobby bonus while waiting for a group to fill, and a USD 0.40 bonus for each correct group decision, averaging about USD 9 per hour.
Measures. Standardised decision accuracy, z-scored across tasks; standardised reliance on correct AI recommendations and on incorrect ones, kept separate; NASA Task Load Index covering mental demand, temporal demand, performance, effort and frustration; perceived teamwork quality, covering the timeliness, precision and usefulness of teammates' contributions; and 5-point evaluations of the devil's advocate on collaboration, satisfaction and quality. Message length was logged as a proxy for deliberation.
Full findings
An LLM that argues against the AI, and keeps arguing, improved group accuracy; an LLM that argues against the group's own majority did not, and the treatment that worked best was the one participants liked least.
Dynamic-AI, the interactive advocate targeting the AI recommendation, was the only treatment to significantly improve group accuracy over control (beta = 0.135, SE = 0.068, 95% CI [0.002, 0.269], p = 0.047). Static-AI marginally raised reliance on correct AI recommendations (beta = 0.200, SE = 0.107, 95% CI [0.010, 0.410], p = 0.062). Both design factors mattered in the expected direction but only marginally: targeting the AI rather than the group majority helped accuracy (F(1,92) = 3.500, p = 0.064, eta-squared = 0.029), and being interactive rather than static reduced over-reliance on incorrect AI recommendations (F(1,91) = 3.554, p = 0.063, eta-squared = 0.038). The direction of the objection is the more novel result: challenging the machine works better than challenging the group, which is the opposite of what classical devil's advocacy research, aimed at group conformity, would predict.
The distribution split (RQ3) shows the two mechanisms are not the same. Dynamic-AI's accuracy gain came mainly from in-distribution cases (beta = 0.122, SE = 0.067, 95% CI [0.009, 0.254], p = 0.069), where the model is 62.5% accurate. Static-AI's gain in correct AI reliance came mainly from out-of-distribution instances (beta = 0.619, SE = 0.355, 95% CI [0.077, 1.315], p = 0.081), the Black-defendant low-prior-crime cases where the model is at 48%. Neither intervention selectively rescued the subpopulation the model was worst on.
The mechanism the authors propose has three parts: interactive advocacy extends deliberation, with Dynamic-AI participants writing significantly longer messages (beta = 198.25, SE = 55.31, 95% CI [89.85, 306.66], p < 0.001); it improves deliberation quality by pushing groups to examine a fuller set of factors and to check the soundness of their arguments; and it works through discomfort. Dynamic-AI participants had the highest accuracy and the lowest self-perceived performance (M = 4.107, SD = 0.791 against Static-AI M = 4.421, SD = 0.617, p < 0.05), and rated teammates' contributions as less precise (M = 3.940 against control M = 4.393, p = 0.010) and less useful (M = 3.821 against control M = 4.410, p < 0.001). Being argued with degraded how the group saw itself while improving what it produced.
Participants nonetheless rated interactive advocates higher than static ones on collaboration (F(1,268) = 9.757, p = 0.002, eta-squared = 0.035) and quality (F(1,268) = 11.211, p = 0.001, eta-squared = 0.040), and the authors note that participants anthropomorphised the advocate, expecting it not only to argue well but to judge when arguing was worth doing at all.