Explanations Can Reduce Overreliance on AI Systems During Decision-Making
N = 731 Logical / reasoning taskThis study in the framework
Fixed at exactly 80%, by construction, and never disclosed to participants. The AI was simulated rather than trained.
The five predictions:
- 1. As tasks increase in difficulty, there will be greater reductions in overreliance with explanations, compared to only getting predictions.
- 2. As the effort to understand the explanation decreases, overreliance decreases.
- 3. As the monetary benefit of correctly completing the task increases, overreliance decreases.
- 4. As tasks increase in difficulty, people will attach higher subjective utility to an AI's explanation.
- 5. As the effort to understand the explanation decreases, people will attach higher subjective utility to an AI.
"people strategically choose whether or not to engage with an AI explanation"
abstract
"some of the null effects found in literature could be due in part to the explanation not sufficiently reducing the costs of verifying the AI's prediction"
abstract
"we have been focused on a narrow slice of the cost-benefit space"
framework
"overreliance is not merely an immutable inevitability of cognition but at least, in part, a strategic choice"
conclusion
"we observe that people will not overrely if they know the AI is wrong"
Study 3 results
Experimental design
Five online experiments on Prolific, N = 731 in total, all using a maze-solving task with a simulated AI.
The task and why it was chosen. Participants see a maze with a start position and four candidate exits and must identify the true exit. The authors set out five requirements and argue the maze meets all of them: it is multi-class, so random guessing yields 25% rather than the 50% a binary task would give; it is non-trivial alone, so AI help is meaningful; it needs no domain knowledge; its difficulty can be tuned continuously by changing dimensions; and it supports multiple explanation modalities. Mazes were generated at 10x10 (easy), 25x25 (medium) and 50x50 (hard). Items were pre-tested with six crowdworkers: easy mazes not solved at 100% accuracy were dropped, and medium and hard mazes solved at 100% were dropped. The authors acknowledge the ecological-validity cost and defend the choice by analogy to game-playing settings such as chess, where a prediction is a move and an explanation is a roll-out.
Explanations. Four variants, all generated to be accurate. Highlight explanations trace the AI's proposed path directly on the maze. Written explanations translate the same path into words ("left, up, right, up...") beside the maze, which raises the cost of use because the participant must parse the text and map it back onto the grid. Two further highlight variants were built to make errors conspicuous: incomplete, which stops drawing the path at the point where the AI crosses a wall, and salient, which continues the path past the wall in blue. When the AI is wrong, its explanation shows a path passing through a wall, so the explanation is always faithful to the prediction.
Study 1 (N = 340). Two-factor mixed design: AI condition (prediction only, or prediction plus highlight explanation) between subjects, task difficulty within subjects. Half of participants saw easy and medium mazes; the other half saw only hard mazes, because the two configurations took similar time. Easy-and-medium participants solved two easy and two medium mazes alone, ten training mazes with the AI including two AI errors, then thirty test mazes with six AI errors at randomised positions. Hard-condition participants did half as many mazes at each phase.
Study 2 (N = 340, pre-registered at osf.io/4dbqp). Same design with explanation modality (highlight against written) as the manipulated factor. 170 new participants in the written condition were compared against the 170 highlight participants from Study 1.
Study 3 (N = 286, exploratory, not pre-registered). Between subjects, hard task only, five conditions: prediction, highlight, written, incomplete, salient. 31 new participants (16 incomplete, 15 salient) were added to 85 each in the prediction, highlight and written conditions carried over from Studies 1 and 2.
Study 4 (N = 114, pre-registered at osf.io/hgz2x). Hard task only. AI condition (prediction or highlight explanation) between subjects; monetary bonus within subjects and blocked, 0.01 USD per correct maze in one half and 0.50 USD in the other. The blocked within-subjects design was chosen because judgements of monetary gain are made relative to a reference point.
Study 5 (N = 76, pre-registered at osf.io/cskvb). Adapts the Cognitive Effort Discounting (COG-ED) paradigm to human-AI teams. Participants repeatedly choose between doing the task with prediction only for a fixed 100 credits and doing it with an explanation for a smaller, dynamically adjusted reward, starting at 50. The reward moves by binary search: it rises toward 100 when the high-effort option is chosen and falls when the low-effort option is chosen. The converged value R6 gives the subjective utility of the explanation, expressed as the reward the participant will forgo to get it. Two within-subjects scenarios: highlight explanations in medium against easy tasks, and highlight against written explanations in the medium task. Credits were used rather than dollars to equalise perceived reward and mitigate income effects.
Measures. Overreliance as the percentage of incorrect AI predictions accepted. Need for Cognition, a stable personality trait, from a six-item scale. Self-reported trust and interaction style on 7-point scales. Subjective utility from COG-ED in Study 5. Accuracy was analysed exploratorily.
Full findings
Explanations do reduce overreliance, but only where the cost-benefit arithmetic favours engaging with the task. The result reframes a decade of null findings as a sampling problem rather than a fact about cognition.
Study 1, task difficulty. Overreliance rose sharply from easy to medium in the prediction-only condition (mean -3.967, 95% CI [-5.480, -2.53], notable), confirming H1a: harder tasks push people toward relying. Explanations made no difference in the easy task (0.164, CI [-1.678, 2.17], H1b confirmed as a null) and none in the medium task (0.916, CI [-0.305, 2.08], H1c not supported). In the hard task, explanations reduced overreliance (1.74, CI [0.889, 2.76], notable, H1d confirmed). An exploratory analysis found explanations also raised decision accuracy in the hard task, which the authors flag as standing against Bansal et al. and Bucinca et al.
The easy and medium nulls are the important part. They replicate the field's standard finding, in the same experiment that overturns it. What changes is only where in the cost space the study sits.
Study 1, Need for Cognition. No interaction between NFC and AI condition in the hard task (0.35, CI [-0.76, 1.49], H1e not supported). The authors suggest the hard task is so demanding that even high-NFC participants overrely. An exploratory analysis in the medium task did find an interaction, with the gap between prediction and explanation narrowing as NFC rose.
Study 2, explanation difficulty. Highlight explanations produced less overreliance than written ones in both the medium task (-1.63, CI [-2.78, -0.559], notable) and the hard task (-1.632, CI [-2.570, -0.694], notable). Both hypotheses supported. An exploratory comparison found no difference between prediction-only and written explanations at any difficulty, which the authors read as evidence that a hard-to-parse explanation does not act as a bare trust signal.
Study 3, pushing saliency. The salient explanation, which highlights the illegal segment of the path in blue, produced an average overreliance rate of 0%. It beat every other condition, including prediction (41.698, CI [5.055, 153.76]) and highlight (39.816, CI [3.843, 152.03]); note the very wide intervals, and that this study was exploratory with only 15 participants in that cell. Incomplete explanations beat prediction (3.336, CI [1.309, 5.36]) and written (3.205, CI [1.236, 5.38]) but did not differ from highlight. The authors conclude there is no floor: people will not agree with an AI they know to be wrong.
Study 4, monetary benefit. Overreliance was lower under the 0.50 USD bonus than the 0.01 USD bonus (0.93, CI [0.414, 1.46], notable, H3a supported), and the hard-task explanation effect replicated with bonuses present (2.47, CI [1.72, 3.36], H3b supported). This is the benefit half of the framework, and the authors note it has a methodological consequence: crowdsourced studies that pay a bonus may not replicate those that do not.
Study 5, subjective utility. Participants assigned higher utility to highlight explanations in the medium task than in the easy task (16.96, CI [5.22, 28.7], notable, H4a) and higher utility to highlight than to written explanations (18.61, CI [9.51, 27.79], notable, H5a). People forgo real money for an explanation in a harder task, or for one that is easier to read. Crucially, the conditions to which they assign the most utility are the same conditions that produced the largest reductions in overreliance, which is what closes the loop between the framework and the behaviour.
The authors reinterpret two rows already in this base. They suggest Bansal et al.'s null arose because sentiment analysis was too easy to complete alone while the LSAT explanations were too complex to check, and that Bucinca et al.'s cognitive forcing functions worked by raising the cost of the relying strategy rather than by improving comprehension.
The practical statement of the result: an explanation reduces overreliance when it lowers the cost of verifying the AI substantially below the cost of doing the task alone. If verification costs about what the task costs, the explanation buys nothing.