Contrastive Explanations That Anticipate Human Misconceptions Can Improve Human Decision-Making Skills
N = 628 Exercise recommendationThis study in the framework
Held constant and simulated rather than trained. The authors set the AI's accuracy at 71.4% (5 of 7) so that it clearly exceeded the roughly 30% unassisted accuracy their formative studies found on a 7-alternative choice, while still erring often enough to study overreliance. Observed unaided accuracy in the study was 0.31, so the AI held a large advantage of roughly 40 points, much larger than in most rows in this base.
The authors group their hypotheses and research questions by outcome, using the suffixes L for learning, A for accuracy and S for subjective experience.
Human learning: H-L1: Contrastive explanations with sensible foil, predicted (H-L1a) or inputted (H-L1b), will lead to more learning than providing people with no AI support. H-L2: Contrastive explanations with sensible foil, predicted (H-L2a) or inputted (H-L2b), will lead to more learning than unilateral explanations. H-L3: Contrastive explanations with predicted foil will lead to more learning than contrastive explanations with a random foil. RQ-L1: Will contrastive explanations with predicted foil, provided at the decision-making time, lead to different learning than contrastive explanations after the decision is made?
Accuracy and overreliance: H-A1: Contrastive explanations with sensible foil, predicted (H-A1a) or inputted (H-A1b), will lead to equal or better decision accuracy compared to unilateral explanations. RQ-A1: Will contrastive explanations, predicted or random, which present two choices, reduce overreliance on AI, compared to unilateral explanations?
Subjective experience: H-S1: Contrastive explanations with predicted foil will lead to higher perceived competence, autonomy, and relatedness to AI than unilateral explanations. H-S2: Contrastive explanations with predicted foil will lead to higher perceived competence, autonomy, and relatedness to AI than contrastive explanations with inputted foil.
"contrastive explanations significantly enhance users' independent decision-making skills compared to unilateral explanations, without sacrificing decision accuracy"
abstract
"contrastive explanations may unevenly impact individuals, offering greater advantages to those with higher AOT"
results, audit for intervention-generated inequalities
"their performance also significantly degraded when AI suggestions were suboptimal compared to receiving no support"
results, accuracy and overreliance
Experimental design
A single between-subjects experiment with five conditions, run on Prolific, N = 628 retained.
Framework. The paper's contribution is a Contrastive Explanation Framework with four components: an AI task model that predicts the AI's answer (the fact); a human model that predicts what an average person would answer for the same item (the foil); a contrast module that identifies the dimensions on which the fact is superior to the foil and, if any, the dimensions on which the foil is superior; and a presentation module, powered by GPT-4, that converts these dimensions into prose and supplies the common-sense knowledge needed to bridge from the model's concept space to the vignette. The argument is that conventional explanations are unilateral, justifying the AI's answer without reference to what the user was thinking, whereas people intuitively seek contrastive explanations.
Task. Exercise recommendation, designed with a kinesiology expert co-author to be accessible to laypeople while imposing the multi-factor weighing typical of clinical treatment selection. Participants read a vignette of a fictitious character and chose the best exercise from a fixed alphabetical list of seven: aerobics, bicycling, boxing, jog/walk combination, pilates, resistance training and swimming. Characters were generated by sampling demographics from US Census, CDC and Bureau of Labor Statistics distributions, then assigning fitness level and maximal intensity, an exercise goal and an exercise preference. Fifty-nine leisure activities were curated from a compendium, labelled by MET intensity, goal (cardio, muscle building, flexibility), indoor or outdoor and individual or group; seven were selected as the answer set. Characters and exercises were encoded into a shared three-concept space of intensity, goal and preference, and the ground truth was the exercise scoring highest under expert model weights learned from the domain expert.
Conditions (five, between-subjects): No AI, the baseline, no support at all. Unilateral, the conventional paradigm: the AI suggests a choice and gives reasoning emphasising all concepts supporting it. Contrastive predicted, in which the explanation contrasts the AI's choice with the foil predicted by the human model, presented in the interface as a choice many people would likely make, and highlights only the concepts on which the two differ. Contrastive random, presentationally identical but with the foil drawn at random from the six alternatives. Contrastive after, in which the participant decides first and their own choice is then used as the foil; where the participant's choice already matched the AI, a unilateral explanation supporting it was shown instead.
Human model. A linear SVM trained on unassisted human responses collected in a separate Prolific study, predicting which of two exercises a lay person would prefer for an unseen character. The foil was the highest-scoring exercise under human weights that was not the expert's choice, that is, the most likely incorrect human answer.
Procedure. Consent, then a demographic survey, a six-item Need for Cognition scale and a seven-item Actively Open-minded Thinking scale. Three blocks followed: a pre-test of 5 unassisted tasks, an intervention block of 14 tasks under the assigned condition, and a post-test of 5 unassisted tasks. Learning is the post-test score controlled for pre-test score. Afterwards participants completed a shortened Intrinsic Motivation Inventory measuring perceived autonomy, competence, relatedness to AI and interest/enjoyment with four items each (three for relatedness), plus a single mental-demand item.
Outcomes and analysis. Accuracy is the proportion correct during the intervention block; overreliance is the proportion of answers matching the AI on items where the AI was wrong; learning is post-test correctness. All three are reported as marginal means from regression models including pre-test performance as a covariate. Learning and accuracy were analysed by ANCOVA with pre-test as covariate and condition as fixed factor, subjective measures by ANOVA, all with Holm-Bonferroni correction across the eight learning-related comparisons. Effect sizes are Cohen's d with 95% confidence intervals; residual normality was checked (Shapiro-Wilk W = .993, p = .137).
Full findings
Contrastive explanations that anticipate what the user was likely to think improved people's unaided decision-making skill without costing accuracy, and conventional unilateral explanations did not. Learning in the contrastive predicted condition reached M = 0.47 against M = 0.32 for no AI (F(1,209) = 38.62, p = .00004, d = 0.65 [0.37, 0.94]), while unilateral explanations at M = 0.39 did not differ significantly from no AI at all (d = 0.30 [0.03, 0.57], n.s. after correction). Directly against unilateral, contrastive predicted learned significantly more (M = 0.47 against 0.39, F(1,260) = 40.99, p = .02, d = 0.35 [0.11, 0.60]), supporting H-L2a.
Accuracy was unaffected in either direction. Contrastive predicted (M = 0.56) and contrastive after (M = 0.57) did not differ from unilateral (M = 0.58), with confidence intervals straddling zero (d = -0.15 [-0.39, 0.10] and -0.08 [-0.32, 0.16]). H-A1a and H-A1b were supported: the learning gain was free.
The contrast in the explanation matters more than the specific foil. Contrastive random also beat no AI on learning (M = 0.41, p = .02, d = 0.40 [0.13, 0.67]), and predicted beat random only marginally (d = 0.26 [0.02, 0.50], p = .09), giving H-L3 partial support. Timing did not matter either: contrastive predicted and contrastive after did not differ on learning (d = 0.18 [-0.07, 0.43]), answering RQ-L1 with a null, and contrastive after did not beat unilateral (d = 0.16 [-0.08, 0.40]), so H-L2b was not supported. What appears to drive the effect is presenting an alternative at all, not tailoring it precisely.
Overreliance was not reduced. Contrastive predicted (M = 0.58) and contrastive random (M = 0.54) matched unilateral (M = 0.58) on the proportion of wrong AI suggestions accepted, and presenting the contrast after the decision made no difference (M = 0.59). RQ-A1 returns a null: showing two options does not make people more discriminating about the one recommended.
The blunt version of the accuracy result deserves stating. AI support raised accuracy overall from 0.31 to 0.56 (F(1,627) = 195.32, p much less than .0001), but on the items where the AI was wrong participants fell to 0.14 against 0.29 unaided (F(1,627) = 36.34, p much less than .0001). Assistance roughly halves unaided performance on the cases the model gets wrong, in every condition including the contrastive ones.