Deciding Fast and Slow: The Role of Cognitive Biases in AI-assisted Decision-making
N = 526 Student performance predictionThis study in the framework
Three different accuracy figures apply. The model. A logistic regression on the UCI Student Performance data, trained on the top 10 features in experiment 1 and deliberately restricted to 7 features in experiment 2 so that participants would hold complementary knowledge. Real accuracy was 70.8% on the training set and 66.5% on the test set, and roughly 60% on the trials used in experiment 2. The stated accuracy. Participants were told at the start of training that the AI is 85% accurate. This was made true over the 15 hand-picked training trials but is well above the model's real performance. The overstatement is deliberate: it is the instrument by which anchoring bias is induced, and the authors defend it as realistic under distribution shift between training and deployment. The displayed accuracy. In experiment 1, eight medium-difficulty trials on which the model was actually correct had their predictions flipped before display, creating probe trials where a knowledgeable participant should disagree. The AI as experienced by participants was therefore 58.3% accurate across the testing section, against a stated 85%.
Experiment 1: H1: Increasing the time allocated to a task alleviates anchoring bias, yielding a higher likelihood of sufficient adjustment away from the AI-generated decision when the decision-maker has the knowledge required to provide a different prediction.
Experiment 2:
- H2: Anchoring bias has a negative effect on human-AI collaborative decision-making accuracy when AI is incorrect.
- H3: If the human decision-maker has complementary knowledge then allocating more time can help them sufficiently adjust away from the AI prediction.
- H4: Confidence-based time allocation yields better performance than Human alone and AI alone.
- H5: Confidence-based time allocation yields better human-AI team performance than constant time and random time allocations.
- H6: Confidence-based time allocation with explanation yields better human-AI team performance than the other conditions.
"the average disagreement percentage in probe trials increased from 48% in the 10-second condition to 67% in the 25-second condition"
results, experiment 1
"the participants in 'Confidence-based time' may have grown to distrust the AI"
discussion, lessons learned
"When the AI is incorrect, the information it provides the human distracts them from the correct decision, thus reducing their performance"
conclusions
Experimental design
Two experiments on Amazon Mechanical Turk, both restricted to US workers with at least a 98% approval rating and 100 approved tasks, using the same task throughout.
Task. Student performance prediction: judge whether a student will pass or fail a class from their characteristics, past performance and demographics. Data from the UCI Student Performance Dataset (Cortez and Silva), 1044 students across mathematics and Portuguese, 33 features. Labels binarised, 70/30 train-test split, logistic regression trained on standardised features, with the top 10 features by coefficient retained: mother's and father's education, mother's and father's jobs, weekly study hours, interest in higher education, weekly hours going out with friends, absences, extra educational support, and past failures.
Inducing the anchor. Both experiments began with a 15-trial training section in which participants predicted first and were then shown the correct answer, plus the AI's prediction, so they could gauge the AI first-hand. Bar charts of outcome distributions per feature were shown during training only, then withdrawn. Training trials were sampled so that the AI's predicted probability was uniform across [0.5, 0.6] through (0.9, 1.0], using predicted probability as a proxy for difficulty. Participants were told the AI is 85% accurate, which was true over the hand-picked training trials but not of the model, whose real accuracy was 70.8% on the training set and 66.5% on the test set. The authors justify the gap as realistic under distribution shift and as necessary to induce anchoring in a short session.
Experiment 1 (N = 47). Testing section of 36 trials, identical and in the same order for all participants, with confidence rated low, medium or high per trial. Two trial types. Probe trials: 8 medium-difficulty trials (predicted probability 0.6 to 0.8) on which the AI was actually correct but its displayed prediction was flipped, so a knowledgeable participant should disagree. Unmodified trials: the remaining 28, sampled to keep predicted probability uniform. Displayed AI accuracy across the testing section was therefore 58.3%, far below the stated 85%.
The time manipulation. The testing section was split into four blocks of 9 trials at 10, 15, 20 and 25 seconds per trial, the sequence being a random permutation per participant, so time allocation is orthogonal to participant, performance and trial order. Each block contained 2 probe and 7 unmodified trials in random order. Participants could not advance to the next trial until the allocated time had elapsed. The four intervals were chosen from a shorter pilot as spanning necessary to sufficient.
Experiment 2 (N = 479). Same task, but complementary knowledge was engineered by training the AI on only 7 features while participants saw 10; the three withheld features (weekly study hours, weekly going-out hours, extra educational support) were the second to fourth most important in a full model. Testing section of 40 trials in randomised order, split by AI confidence at a threshold of 0.75 into 20 low-confidence and 20 high-confidence trials.
Five between-subjects groups: Human only (no AI, 25 seconds per trial, n = 95); Constant time (AI, 17.5 seconds reported as 18, n = 109); Random time (AI, drawn uniformly from 10 or 25 seconds, mean 17.5, n = 95); Confidence-based time (AI, 25 seconds on low-confidence trials and 10 on high-confidence, n = 85); and Confidence-based time with explanation (as above, plus the AI's confidence displayed as low or high, n = 96). Time allocations were grouped into 8 blocks of 5 trials to avoid rapid switching.
Payment. Experiment 1: $3.50 base plus $1 accuracy bonus, mean 27 minutes, about $10 per hour. Experiment 2: mean $4.125 base plus $1 bonus, mean 30 minutes, about $10.25 per hour.
Full findings
Time reduces anchoring, but only if you tell people why they are being given it.
Experiment 1 established the basic effect. On probe trials, where the displayed AI prediction had been flipped and participants had the knowledge to catch it, average disagreement rose from 48% at 10 seconds to 67% at 25 seconds. A bootstrap regression of disagreement on allocated time gave a significantly positive coefficient of 0.01 (95% CI [0.001, 0.018]), corresponding to a 0.15 rise in disagreement between the 10-second and 25-second conditions. H1 was supported. The authors note this validates empirically an idea proposed by Tversky and Kahneman but not previously tested in the AI-assisted setting: insufficient adjustment away from an anchor is a resource-rational trade-off between accuracy and time, and buying more time buys more adjustment.
Disagreement on unmodified trials stayed near 0.1 at every time allocation, which the authors read as showing participants had roughly the AI's level of knowledge and nothing extra to contribute when the AI was right. That motivated Experiment 2's engineering of complementary knowledge.
Experiment 2 confirmed the harm and complicated the remedy. The setup worked: AI alone and Human only reached similar overall accuracy near 60%, yet participants agreed with the AI on only 62.3% of trials, and Human only accuracy on AI-incorrect trials was 61.8% in low-confidence trials against 29.8% in high-confidence ones, so complementary knowledge was concentrated where the AI was unsure.
Anchoring damaged the team. Agreement was much higher in the collaborative conditions than in Human only (p < .001, t(184) = 6.73), and when the AI was wrong this cost accuracy relative to Human only (p < .001, t(370) = -6.68). Because Human only had more time on average, the authors repeat the comparison at matched time, Human only against Confidence-based time on low-confidence trials at 25 seconds each, and find the same disparity in agreement (p < .001, t(92) = 4.97) and accuracy (p < .001, t(92) = -4.74). H2 supported.
The key result is that time allocation alone was not enough. Confidence-based time by itself did not produce sufficient adjustment away from incorrect AI predictions. Adding the confidence disclosure did: Confidence-based time with explanation significantly reduced anchoring on low-confidence trials against all other conditions (p = .003, t(383) = 2.70), lifting accuracy on those trials to 43.8% against 37.5% for Confidence-based time alone, 36.4% Constant and 36.2% Random. H3 was supported only in the explanation arm.
Overall team accuracy barely moved: 61.9% for Confidence-based with explanation, 61.5% Constant, 61.1% Random, 61.0% Confidence-based. Both confidence-based conditions beat AI alone (p = .06 and p = .004) but neither beat Human only, so H4 and H6 receive only partial support, and H5 none.