← Back to the framework

The Amplifying Effect of Explainability in AI-assisted Decision-making in Groups

N = 89 Mushroom classification

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Manipulated
Composition Manipulated (Individual, Group)
Size 2
Mode Interdependent (deliberating)
Task Inherited
Difficulty Not reported
Stakes Low
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 80%-90%
AI design
Protocol Human-first (update)
AI stance Directive
Interactivity Static
XAI Manipulated
Presence Manipulated
Type None · Feature importance
Quality Correct
Number of AI advisors One
Research questions
  • RQ1: In AI-assisted decision-making scenarios, what impact do AI explanations (XAI) have on the performance of decision-making by two-person teams?
  • RQ2: How reliance on AI differs between decisions made by individuals alone versus decisions made by two people working together as a pair?
In the authors' words

"Groups rely less on incorrect AI recommendations when explanations are available, but they rely more on incorrect AI recommendations when explanations are absent, compared to individual decision makers."

abstract

"In group settings, this effect is magnified, as the AI recommendation becomes an anchor for multiple participants."

discussion, 5.2

"Explanations encouraged groups to transition from fast, heuristic decision-making to a slower, more reflective approach, enabling members to collaboratively clarify misunderstandings or discuss misalignments in their decisions. This dynamic is not accessible to individuals working alone, even when explanations are provided."

discussion, 5.2

"our findings strongly indicate that reliance behavior does not align with reported trust, at least within the context of one-shot studies."

discussion, 5.1
Experimental design

A 2x2 between-subjects design crossing the presence of explanations (XAI) with the number of human decision-makers (Team Composition), run in person.

Task. Classifying a mushroom as edible or poisonous from a visual representation and eight categorical characteristics: odour, cap colour, cap shape, cap surface, gill colour, gill size, gill spacing and ring number, with between 2 and 12 levels each. The task was chosen for two stated reasons: it carries a psychological element of risk, though not a real one, since misclassifying a poisonous mushroom feels consequential; and it is intelligible without mycological knowledge, so it accommodates varied participant expertise.

Procedure per round. Four steps. The participant or pair judged edibility from the characteristics alone, then rated confidence in that judgement from 0 to 10, then saw the AI recommendation, with or without an explanation depending on condition, and made a second judgement, then received feedback on the mushroom's true edibility. Requiring an independent first judgement is what makes the reliance decomposition possible.

Structure. A two-round tutorial, then 15 rounds of the main task. The order of the 15 mushrooms was identical for everyone, and the AI was correct on 12 of them, with errors fixed at positions 5, 10 and 13, giving a perceived accuracy of 80%. Across 59 sessions this yields 885 decisions.

Conditions. Team Composition: Individual, one participant deciding alone; Group, two participants working collaboratively and encouraged to discuss and reach consensus before each decision. Pairs volunteered together, so they already knew each other and shared backgrounds. XAI: Without XAI, only mushroom characteristics and the AI prediction; With XAI, the same plus an explanation showing the importance of each characteristic to the prediction, generated with LIME. Feature importance was chosen deliberately over counterfactual or example-based explanations on the grounds that lay users lack any mycological frame of reference.

AI system. A decision tree trained on a tabular mushroom dataset, with 96% accuracy on its test set. The 80% figure participants experienced was imposed by the fixed error positions rather than sampled.

Measures. Behavioural: decision accuracy for both the initial and the final decision; decision confidence, 0 to 10; decision time for both decisions; and reliance, over-reliance and under-reliance rates following Schemmer et al., each computed within its own domain. A case enters the reliance domain only when the initial human decision differs from the AI prediction; reliance is then the fraction of those cases where the final decision matches the AI. Over-reliance is restricted to the subset where the AI was wrong, under-reliance to the subset where the AI was right and the participant kept their own answer. Subjective, collected individually even in the group condition: trust on the Multi-Dimensional Measure of Trust across eight traits (reliable, predictable, consistent, skilled, capable, competent, precise, transparent), 0 to 7; understandability on three items, 0 to 7; plus two open-ended questions on reasoning and on how AI recommendations were used.

Full findings

Explanations did not change what groups did on average; they changed how far groups swung. Adding a human partner amplified the effect of explanations on over-reliance in both directions, so the same manipulation that made pairs the most careful decision-makers in the study made their absence the most careless.

The interaction is the result. Over-reliance showed a main effect of XAI, F(1, 55) = 9.523, p = 0.004, eta-squared = 0.192, and an interaction with team composition, F(1, 55) = 6.098, p = 0.018, eta-squared = 0.132, with no main effect of team composition at all (F = 0.006, p = 0.94). Within the group condition, over-reliance was 0.82 (SD = 0.40) without explanations and 0.125 (SD = 0.23) with them; the simple effect of XAI is significant only in the group condition, p < 0.001. Groups with explanations also over-relied less than individuals with explanations (0.125 against 0.42, SD = 0.49). A pair given a bare recommendation followed a wrong AI in four cases out of five; the same pair given a feature-importance chart followed it in one case out of eight.

Under-reliance did not move for anyone: no effect of XAI (F = 0.256, p = 0.615), of team composition (F = 0.011, p = 0.917) or of their interaction (F = 1.720, p = 0.195). Overall reliance showed only a marginal interaction, F(1, 55) = 3.97, p = 0.05, which does not survive the Bonferroni threshold of 0.016. The authors are careful about this and say so. The asymmetry is their substantive claim: explanations fix the failure of following a wrong system and do nothing about the failure of ignoring a right one, which they attribute to algorithm aversion and overconfidence being more robust biases than automation bias.

Accuracy was affected by explanations and not by team composition. Final decision accuracy was 0.76 (SD = 0.094) with explanations against 0.66 (SD = 0.142) without, F(1, 55) = 9.98, p = 0.003, while groups and individuals were indistinguishable at 0.711 (SD = 0.143) and 0.710 (SD = 0.116), F(1, 55) = 0.019, p = 0.892. AI assistance helped overall: final accuracy 0.71 (SD = 0.13) against initial 0.60 (SD = 0.14), F(1, 55) = 41.2, p < 0.001. Initial decision accuracy also differed by XAI condition, 0.65 against 0.55, F(1, 55) = 8.268, p = 0.005, which is a striking result for a judgement made before any AI output is shown and is discussed in Observations.

The proposed mechanism is deliberation time. Over-reliance correlated negatively with time spent on the final decision within groups, r(30) = -0.65, p < 0.001, while reliance did not, r(30) = -0.29, p = 0.12, and under-reliance did not, r(30) = 0.06, p = 0.75. Explanations lengthened the final decision, 13.90 seconds (SD = 8.96) against 7.68 (SD = 5.24), F(1, 55) = 9.97, p = 0.003, and groups took longer on the initial decision than individuals, 27.86 (SD = 11.83) against 19.66 (SD = 9.67), F(1, 55) = 8.502, p = 0.005. The authors read the group amplification as anchoring and automation bias magnified by having several people anchored on the same recommendation, plus diffusion of responsibility, plus a groupthink route in which one member's rationalisation carries the pair; explanations counteract this by giving the pair something concrete to talk about. The qualitative analysis supports it: 50% of paired participants with explanations reported using the AI to learn the task against 27% of individuals with explanations, and participants described the AI as a third group member, one saying they treated it like another human and that in a democracy the majority wins.