Do People Engage Cognitively with AI? Impact of AI Assistance on Incidental Learning
N = 1010 Nutrition estimationThis study in the framework
AI Accuracy was 100% by construction, and deliberately so. The Meal Assistant was simulated rather than trained, and its recommendations were always correct in every condition of every experiment. Participants were nonetheless told that it is right most of the time but not always, language the authors chose because that uncertainty is typical of contemporary decision support and can shift trust.
RQ: Do people engage cognitively with AI-generated assistance deeply enough for incidental learning to occur?
Experiment 1 hypothesis: a simple explainable AI design, offering a decision recommendation and an explanation, would improve immediate task performance compared to no AI support, but would not result in learning.
Experiment 2 hypothesis: the Update design, in which the person decides first and only then sees the AI recommendation and explanation, would result both in improved task performance and in learning. This followed prior work showing that the Update design improves accuracy and reduces overreliance, and it induces cognitive engagement.
Experiment 3 hypothesis: an AI explanation only design, providing the information needed to decide but no explicit recommendation, would result both in improved task performance and in learning. This followed research on navigation aids showing that withholding an explicit suggestion prompts active processing.
"we designed our study such that the Meal Assistant recommendations were always correct"
methods, experiment 1
"Contrary to our expectations, participants did not learn more in the Update condition than in the Minimal feedback condition"
results, experiment 2
"in the AI explanation only condition the High NFC group learned significantly more than the Low NFC group"
results, audit for intervention generated inequalities
Experimental design
Three within-subjects experiments plus an independent replication of the third, on the same task with the same two non-AI baselines throughout.
Task. A nutrition knowledge quiz. Each question shows images and descriptions of two meals and asks which contains more of one macronutrient (carbohydrates, fat, protein or fibre). Every question turns on a single nutrition concept, for example that avocados are a significant source of fat or that soy beans contain more protein than most other beans. The authors stress that this is a real decision task rather than a proxy task.
Learning design. For each concept three questions were prepared, serving as pre-test, intervention and post-test. The pre-test measures pre-existing knowledge and returns correctness feedback only, a choice justified by prior work showing that bare correctness feedback produces about as little learning as no feedback, while pilot participants preferred receiving something. The intervention question is where the experimental condition applies. The post-test measures what was learned from the intervention, and afterwards participants received both correctness feedback and an explanation purely to improve their experience. Each study drew eight concepts, two per macronutrient, giving 24 questions. Concepts were randomly assigned to conditions and question order was randomised, so different questions served as pre-test, intervention and post-test for different participants.
Conditions. Two baselines ran in every experiment. Minimal feedback (low baseline): no AI, correctness feedback only after answering. Prior work found this indistinguishable from no feedback at all, so little else could produce less learning. Explanation feedback (high baseline): no AI, correctness feedback plus a brief explanation after answering, given regardless of whether the answer was right. Prior work found this produces significantly more learning than correctness feedback alone, so it represents a demonstrably effective design.
The third condition varied by experiment. Experiment 1, AI recommendation and explanation: a simulated Meal Assistant gives both a recommendation and an explanation at the moment the question is posed, before the participant decides. No feedback follows. Experiment 2, Update: the participant answers first, then sees the Meal Assistant's recommendation and explanation and may revise. The assistant appeared regardless of whether the initial answer was correct, and participants did not know in advance which trials would carry it. Experiment 3, AI explanation only: the explanation appears before the decision but no recommendation does, so the participant must apply the information themselves. No feedback follows.
The AI. Introduced as an experimental computer system that can analyse the nutritional content of meals, and described to participants as right most of the time but not always. In fact its recommendations were always correct, a deliberate choice because the study targets cognitive engagement rather than trust. Explanations were designed as contrastive explanations, with the foil implicit because each question is binary, and were identical to those used in the Explanation feedback baseline.
Measures. Immediate benefit is the normalised change in correct rate from pre-test to intervention; learning is the normalised change from pre-test to post-test, using the standard normalised change formula from the education literature. Experiments 2 and 3 added a four-item abbreviated Need for Cognition scale.
Full findings
The recommendation is what destroys learning, not the absence of explanation. Across three experiments the same explanation content produced learning or not depending entirely on whether an explicit recommendation accompanied it.
Experiment 1 confirmed the authors' pessimistic hypothesis. Showing an AI recommendation with an explanation produced the largest immediate benefit (M = 0.341) against both the Explanation feedback baseline (M = 0.158, Z = 3.55, p = .0004, r = 0.22) and the Minimal feedback baseline (M = 0.167, Z = 3.56, p = .0004, r = 0.22). But learning was another matter: the Explanation feedback baseline produced more learning (M = 0.325) than the AI condition (M = 0.206, Z = 2.37, p = .0179, r = 0.15), and the AI condition was statistically indistinguishable from minimal correctness feedback (M = 0.189, Z = 0.68, n.s.). People performed better while learning no more than if they had been told only whether they were right.
Experiment 2 delivered the paper's surprise. The Update design, which prior work credits with improving accuracy and reducing overreliance and which Bucinca et al. proposed induces cognitive engagement, improved performance substantially: final answers reached M = 0.391 against Minimal feedback at M = 0.110 (Z = 5.03, p < .0001, r = 0.31) and Explanation feedback at M = 0.178 (Z = 3.67, p = .0002, r = 0.22), with the gain coming demonstrably from revision, since final answers beat participants' own initial ones (Z = 8.17, p < .0001, r = 0.50). Yet learning in the Update condition (M = 0.121) did not exceed minimal correctness feedback (M = 0.157, Z = 0.71, n.s., r = 0.04), and the Explanation feedback baseline again beat it (M = 0.354, Z = 4.59, p < .0001, r = 0.28). Committing to an answer before seeing the AI is enough to reduce overreliance but not enough to produce learning.
Experiment 3 identified what does work. Withholding the recommendation while keeping the explanation produced the largest immediate benefit of any condition in the paper (M = 0.422 against Minimal feedback M = 0.158, Z = 4.54, p < .0001, r = 0.31, and Explanation feedback M = 0.144, Z = 4.84, p < .0001, r = 0.33) and, unlike every other AI condition, real learning (M = 0.342 against Minimal feedback M = 0.138, Z = 3.39, p = .0007, r = 0.23), statistically indistinguishable from the high baseline (Z = 0.41, n.s., r = 0.03). An independent replication on 270 fresh participants reproduced every significant difference at comparable effect sizes.
Read together the three experiments isolate the mechanism cleanly, because explanation content was held identical throughout. When a recommendation is present the person can reach the right answer without processing the explanation, and does. Remove the recommendation and the explanation becomes the only route to the answer, so it must be processed, and processing it leaves knowledge behind. Performance and learning are not the same outcome and can be traded against each other by interface design alone.
The internal audit qualifies the recommendation. Splitting participants at the Need for Cognition median, the two groups did not differ in either baseline condition (Z = 0.668 and Z = 0.677, both n.s.), but in the AI explanation only condition high-NFC participants learned significantly more (M = 0.388) than low-NFC participants (M = 0.255, Z = 2.35, p = .02). The intervention that produces learning produces more of it for people already disposed to think hard, so deploying it risks widening a gap rather than closing one.