To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI
N = 199 Nutrition estimationThis study in the framework
Held constant. A simulated AI with 75% accuracy at identifying the ingredient highest in carbohydrates. No model was trained; the authors state they wanted control over the type and prevalence of error.
H1a: compared to simple explainable AI approaches, cognitive forcing functions will improve the performance of human+AI teams in situations where the AI's top prediction is incorrect.
H1b: compared to simple explainable AI approaches, cognitive forcing functions will improve the performance of human+AI teams.
H2: there will be a negative correlation between the self-reported acceptability of the interface and the performance of human+AI teams in situations where the AI's prediction is incorrect.
"people rarely engage analytically with each individual AI recommendation and explanation, and instead develop general heuristics about whether and when to follow the AI suggestions"
abstract
"there was a trade-off: people assigned the least favorable subjective ratings to the designs that reduced the overreliance the most"
abstract
"explanations are interpreted as a general signal of competence—rather than being evaluated individually for their content—and just by their presence can increase the trust in and overreliance on the AI"
introduction
"our research suggests that human cognitive motivation moderates the effectiveness of explainable AI solutions"
discussion
Experimental design
A mixed between- and within-subjects experiment on a nutrition task, contrasting three cognitive forcing designs against two simple explainable AI designs and a no-AI baseline.
Task. Participants were shown images of meals and asked to replace the ingredient highest in carbohydrates with one low in carbohydrates but similar in flavour. Each decision had two discrete parts: selecting the ingredient to remove from a list of roughly ten options that included all main ingredients in the image plus distractors, then selecting the replacement from a second list of roughly ten. Every meal was constructed so that a single ingredient carried substantially more carbohydrates than the rest. Nutrition was chosen deliberately as a domain approachable to laypeople, the authors noting that access to judges or clinicians is prohibitively costly.
Conditions, in the authors' names. The no AI condition showed the meal image and the two pull-down menus alone. Two simple explainable AI baselines: explanation showed the recognised ingredients and the top four substitutions immediately, each substitution annotated with its estimated carbohydrate reduction and flavour similarity; uncertainty was identical but added the prompt "the AI is X % confident in its suggestion", displayed only on those items where the AI was uncertain. Three cognitive forcing functions: on demand withheld the suggestion until the participant clicked "see AI's suggestion"; update required an unassisted decision first, after which the suggestion and explanation appeared and the decision could be revised; wait displayed "AI is processing the image" for 30 seconds before revealing the suggestion. The explanation content was identical across all five AI conditions.
Procedure. 26 questions in two blocks of 13, the first question of each block a discarded practice item, leaving 24 analysed. Six items carried incorrect model predictions, three per block, at fixed positions; which items were correct or incorrect was randomised in order within that constraint. Each participant saw a different condition in each block, drawn at random from nine conditions in total, of which six are reported here and three were exploratory. A subjective questionnaire followed each block; the four-item need for cognition scale was administered after the first block.
Measures. Objective: overall performance as the percentage of top replacements, meaning both the correct ingredient removed and the top substitution chosen; carb source detection as the percentage of correct removals; carbohydrate reduction achieved; flavour similarity achieved; over-reliance as the percentage of agreement with the AI on items where it was wrong; and human error as the percentage of incorrect decisions that also differed from the AI's suggestion. Subjective, all on 5-point Likert items after each block: preference ("I would like to use this system frequently"), trust ("I trust this AI's suggestions for optimal replacement"), mental demand ("I found this task difficult") and system complexity ("the system was complex").
Full findings
Cognitive forcing reduced over-reliance but did not raise overall team performance, and the designs that protected people best were the ones they liked least. The paper's most consequential result is that this trade-off is not incidental: across participants, performance on items where the AI was wrong correlated negatively with trust and preference.
On items where the AI erred, the ordering was simple explainable AI worse than cognitive forcing worse than no AI at all. Overall performance was 0.03 under simple explainable AI, 0.09 under cognitive forcing and 0.18 with no AI (F(2,261.7) = 12.59, p much less than .0001, d = .37). Carb source detection was 0.08, 0.27 and 0.49 respectively (F(2,242.6) = 35.59, p much less than .0001, d = .66). Unassisted participants working the same items outperformed both assisted groups. Over-reliance on carb source detection fell from 0.64 under simple explainable AI to 0.48 under cognitive forcing (F(1,145.8) = 9.24, p = .003, d = .36), and the distribution of over-reliance, human error and correct decisions differed significantly between categories (chi-square(2, N = 663) = 44.35, p much less than .0001). Human error did not differ between categories on any measure, so the gain came from converting over-reliance into correct decisions rather than from making people generally more careful.
H1a was supported and H1b was not. Pooled across all items, correct and incorrect, both assisted categories beat the no-AI baseline (overall performance 0.35 simple explainable AI, 0.33 cognitive forcing, 0.17 no AI; F(2,141.6) = 15.43, p much less than .0001), but cognitive forcing and simple explainable AI did not differ from one another on any objective measure, and no individual design differed from another within its category. The gain on error items was therefore offset elsewhere. Both assisted categories also remained below the AI's own 75% accuracy.
The acceptability trade-off, H2, held. Cognitive forcing was rated more complex than simple explainable AI (system complexity 2.95 against 2.64; F(2,172.2) = 6.19, p = .002) and trusted slightly less (3.72 against 3.91, n.s.). Across participants, trust correlated negatively with overall performance on incorrect predictions (r = -.24, p = .0003) and with carb source detection on those items (r = -.29, p = .0008); preference likewise (r = -.16, p = .0002 and r = -.21, p = .0005). Mental demand correlated positively with carb source detection on incorrect predictions (r = .13, p = .03) while correlating negatively with performance on correct ones. Trust also predicted over-reliance directly: it correlated with reliance on incorrect predictions for carb source detection (r = .14, p = .04) but not with reliance on correct predictions, which is the signature of miscalibration rather than of well-founded confidence.