Cognitive Forcing for Better Decision-Making: Reducing Overreliance on AI Systems Through Partial Explanations
N = 474 Logical / reasoning taskThis study in the framework
Held constant at a simulated 75% in both studies, and simulated in the strict sense: there was no model. Nine of the twelve suggestions per study were correct and three were incorrect, with the incorrect one distributed one per difficulty level so that accuracy did not covary with difficulty.
- RQ1: Does a partial explanation, revealing only a fragment of the AI's reasoning, reduce overreliance on incorrect AI suggestions relative to a full explanation and relative to no explanation at all?
- RQ2: How does task difficulty moderate the effect of explanation style on decision accuracy and on reliance?
- RQ3: Do individual differences in need for cognition, and in subjective numeracy or reading proficiency, change how much a participant benefits from an explanation?
"Our results show that partial explanations reduce overreliance on incorrect AI suggestions, performing significantly better than the baseline but not as well as full explanations."
abstract
"This means that partial explanations result in significantly less overreliance than full explanations but also introduce underreliance on correct AI suggestions."
results, Study I, AI correctness
"These differences pose a challenge for the design of decision-making systems, as ways of engaging people with AI-assisted tasks disproportionately benefit individuals with high NFC or reading proficiency."
discussion
Experimental design
Two separate between-subjects experiments, each crossing explanation style between participants with task difficulty and AI correctness within participants. The two studies use independent samples drawn from the same pool under identical recruitment criteria. Study I task. Twelve weighted bidirectional graphs; participants entered the cost of the shortest path between a given start and end node. A free numeric response was chosen deliberately over a binary choice, because with a binary task random guessing overlaps heavily with the AI suggestion and contaminates the reliance measure. Participants gave 8.33 distinct answers per task on average, which the authors report as confirmation that the measure works.
Study I conditions. Four explanation styles assigned between participants: baseline, the suggested path cost with no explanation (n = 66); single edge, highlighting one randomly chosen edge of the suggested path (n = 64); ends, highlighting the first and last edges of the suggested path (n = 67); full, highlighting the entire suggested path (n = 67). Single edge and ends are the two partial explanations.
Study II task. Twelve short text segments drawn from the Cambridge English Write & Improve corpus, presented as screenshots to stop participants pasting the text into a spellchecker. Participants entered the number of spelling and grammar mistakes. They were told that every mistake is a one-to-one word replacement and to disregard capitalisation and punctuation.
Study II conditions. Three explanation styles: baseline, the suggested mistake count alone (n = 70); partial, underlining the words the system flagged (n = 78); full, underlining the flagged words and printing a suggested replacement above each (n = 62).
Operationalising difficulty. Study I used graph size: 6 nodes for easy, 10 for medium, 14 for hard, with five graphs at each level. Study II used the Flesch-Kincaid reading ease test, with easy segments at a mean score of 80.2 (SD = 4.92), medium at 60.8 (SD = 4.44) and hard at 38 (SD = 2.35). Difficulty varied within participants, and the ordering was identical for everyone so that order could not confound the result. Self-reported difficulty was collected as a manipulation check: it rose monotonically in Study I, 1.90, 2.99 and 3.09 on a five-point scale, but barely moved in Study II, 2.62, 2.72 and 2.83.
AI system. There was no model. The suggestions were authored, and participants were led to believe an AI had produced them. In both studies 9 of the 12 suggestions were correct and 3 incorrect, one per difficulty level, simulating 75% accuracy. Participants were told the AI was imperfect but never told the accuracy figure, and were never shown the correct answer, because seeing a system err changes how it is perceived. In Study II half the incorrect suggestions were false positives, flagging a correct word and inflating the count by one, and half were false negatives, omitting a required replacement and deflating the count by one.
Measures. Per task: performance as a binary match against ground truth; reliance as a binary match between the participant's answer and the AI suggestion; and five-point ratings of perceived task difficulty, perceived usefulness of the AI and perceived confidence in the AI. After the tasks: the eighteen-item need for cognition scale, five-point trust in the AI, five-point satisfaction with the presentation of the explanations, and the eight-item subjective numeracy scale in Study I, replaced in Study II by a five-point self-rating of reading proficiency.
Full findings
Partial explanations buy a reduction in overreliance and pay for it in underreliance, and the net effect on accuracy is negative relative to a full explanation.
Explanations of any kind beat no explanation. Study I performance rose from 56% (SD = 26%) in the baseline to 70% (SD = 24%) overall, and Study II from 29% (SD = 17%) to 43% (SD = 22%). In the Study I performance model both partial conditions improved on the baseline (single edge estimate 0.63, p = 0.011; ends 0.73, p = 0.003) and the full explanation improved on it most (1.23, p < 0.001); after Tukey correction the ends and full contrasts against baseline survived (p = 0.016 and p < 0.001) while the single-edge contrast fell just short (p = 0.053). Study II reproduced the ordering with every contrast significant: partial over baseline (p = 0.013), full over baseline (p < 0.001), and, critically, full over partial (p = 0.001). The full-over-partial gap was not significant in Study I (single edge p = 0.086, ends p = 0.192), so the accuracy cost of partiality is established in one study of the two.
The reliance models carry the mechanism, and there the result is unambiguous in both studies. Every explanation condition raised reliance over the baseline (all p < 0.001), and both partial conditions produced significantly less reliance than the full explanation (Study I p < 0.001; Study II p < 0.001), with the two Study I partial variants indistinguishable from one another. Splitting reliance by AI correctness, partial explanations reduced reliance on correct and on incorrect suggestions alike (all p < 0.001). The authors state the trade-off plainly: partial explanations result in significantly less overreliance than full explanations but also introduce underreliance on correct AI suggestions. Nothing in the design allows the reduction in overreliance to be obtained without the corresponding loss on correct advice.
The difficulty result is the paper's unexpected finding, and it is non-monotonic. In both studies participants performed worse on medium tasks than on hard ones (Study I p = 0.005; Study II p < 0.001), against the obvious prediction. Study I supplies the reliance evidence for the authors' explanation: reliance was significantly lower on medium tasks than on easy ones (p < 0.001), while reliance on hard tasks was statistically indistinguishable from easy (p = 0.976). The reading is that a visibly daunting graph pushes people to defer to the AI, which at 75% accuracy beats their own attempt, whereas a medium graph looks tractable and invites an unsuccessful attempt. Study II did not replicate the reliance pattern, no difficulty effect on reliance at all, which the authors attribute to text difficulty not being legible at a glance the way graph size is. The performance discrepancy therefore replicated without its proposed mechanism.
Need for cognition moderated overreliance rather than reliance in general. In Study I it was uncorrelated with performance overall (p = 0.126) and with reliance overall (p = 0.677), but correlated positively with performance on tasks where the AI was wrong (r(262) = 0.15, p = 0.018) and negatively with reliance on incorrect suggestions (r(262) = -0.22, p < 0.001). Study II inverted the shape of the benefit: high need for cognition and high reading proficiency correlated with reliance on correct suggestions (r(208) = 0.14, p = 0.042 and r(208) = 0.18, p = 0.010) and with the resulting score. The authors draw a methodological point that transfers beyond this paper: student samples are likely to be high on need for cognition, so a literature recruiting students will overestimate how well cognitive-forcing interventions work.
Attitudes ran against behaviour. Participants were most satisfied with the full explanations, the ones that demanded least of them, and satisfaction correlated with confidence in the AI and negatively with perceived difficulty in both studies. Perceived difficulty was significantly lower for incorrect AI suggestions than for correct ones in Study I (F(1) = 9.50, p = 0.002), which the authors flag as surprising and leave unresolved.