← Back to the framework

Accuracy-Time Tradeoffs in AI-Assisted Decision Making under Time Pressure

N = 475 Logical / reasoning task

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Manipulated
LowHigh
Stakes Not reported
Stress level / time constraint Manipulated
LowHigh
Task uncertainty Not reported
AI Ecosystem Designable
AI performance 70%-80%
AI design Manipulated
Protocol Manipulated (AI-first, Human-first (update), No AI (control))
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Feature importance
Quality Correct
Number of AI advisors One
Research questions

Experiment 1 and repeated in experiment 2: H1: Participants use different AI assistance types differently under time pressure compared to under no time pressure. H1.1: On AI-before, people overrely more under time pressure. H2: There is some underlying trait that can predict whether or not a person overrelies more on AI assistance. H2.1: We can predict whether a person overrelies more or not in the second half of the study from their overreliance behaviour in the first half.

Experiment 2 only: H3: In the mixed condition, overreliance rate on AI-before is higher than for the pure AI-before condition. H3.1: Overreliance rate is higher under no time pressure. H3.2: Overreliance rate is higher under time pressure. RQ1: Can we use personality traits to predict whether or not a person overrelies more on AI assistance?

A third section is exploratory: how AI assistance might be adapted to properties of the person (whether they overrely) and of the task (whether the question is easy or hard).

In the authors' words

"time pressure affects how users use different AI assistances, making some assistances more beneficial than others when compared to no-time-pressure settings"

abstract

"the scarcity effect is not additive with time pressure: time pressure already increases overreliance, and the scarcity effect does not increase this further"

results, experiment 2

"the not-overrelier group achieved human-AI complementarity, while the overrelier group did not"

introduction, summarising the exploratory analysis
Experimental design

Two experiments on Prolific, restricted to English speakers in the United States, sharing one task and one AI.

Task. An alien prescription task adapted from Lage et al. Participants act as doctors treating a series of sick aliens across timed medical shifts, prescribing one medicine per alien. Each alien comes with observed symptoms and a unique treatment plan expressed as a decision set of rules. Participants must derive the medicine using only the observed symptoms and any intermediate symptoms these imply. The authors extended the original setup in three ways: intermediate symptoms were always present, requiring an extra deductive step and providing the material for the AI's explanation; two correct medicines existed per alien, with the better one addressing more observed symptoms, so that a suboptimal but verifiably correct recommendation could be over-relied on more readily than an outright wrong one; and questions came in two designed difficulty levels, easy and hard, constructed to look superficially alike in line count and length so that difficulty was not visually detectable.

AI assistance. When shown, a red box gave both a recommendation and an explanation, the explanation always being an intermediate symptom leading to the recommended medicine. Four assistance conditions: No-AI, with no assistance; AI-before, with recommendation and explanation shown alongside the question before any decision; AI-after, the update protocol, in which the participant decides first and may then revise after seeing the recommendation; and Mixed, in which each individual question randomly received one of the three. Half as many participants were allocated to No-AI as to each other condition, since it served only as a complementarity baseline.

Time pressure. Implemented as two on-screen timers. A global timer counted down the shift, and a local timer reset to 60 seconds per question as a recommended time, turning orange at 10 seconds and red at 0 before running negative. In AI-after, the local timer was reset to 20 seconds after the initial response. Nothing was enforced; the timers exerted pressure without imposing a deadline. The local timer was included specifically so that pressure applied evenly across the shift rather than building towards the end.

Experiment 1 (N = 159 of 207 recruited). Mixed design: AI condition between-subjects, time pressure within-subjects. Four shifts of 5 minutes with break screens, alternating timer and no-timer, with the starting state randomised.

Experiment 2 (N = 316 of 403 recruited). Fully between-subjects, 4 AI conditions by 2 time-pressure levels, giving 8 cells, in a single 20-minute block. This was designed to remove the carry-over that experiment 1 revealed, and to match the structure of the existing literature, where participants are typically never under time pressure.

Procedure (identical in both). Consent, then three survey pages covering demographics, four Need for Cognition items and two self-reported time-pressure-performance items, and Big-5 traits via BFI-10 plus two TIPI neuroticism items. Then instructions and three practice questions with feedback, with a second set allowed on failure. No feedback was given during the main study. Roughly half the questions were easy, each having an independent 50% chance. An exit survey asked perceived task difficulty and perceived AI helpfulness on 5-point scales plus three open questions on strategy.

Outcomes. Accuracy scored 1 for the best medicine, 0.5 for a suboptimal but correct one and 0 for a wrong one. Response time per question. Overreliance defined as the proportion of times the participant gave the same answer as the AI when the AI was wrong or suboptimal. Analysis used ANOVA with Tukey HSD post-hoc, Holm-Bonferroni correction for the within-subject and multi-condition comparisons, logistic regression with chi-squared tests for predicting overreliance group, and Pearson correlations for the personality analyses.

Full findings

Time pressure changes which assistance protocol is best, and it does so by moving the accuracy-time tradeoff rather than accuracy alone. In experiment 2, under no time pressure the three AI-assisted conditions were statistically indistinguishable on accuracy (F(2,135) = 1.46, p = .24), response time (F(2,135) = 0.15, p = .86) and overreliance (F(2,135) = 2.37, p = .10). Under time pressure, accuracy stayed indistinguishable (F(2,135) = 0.80, p = .45) but response time (F(2,135) = 5.09, p = .0073) and overreliance (F(2,135) = 4.09, p = .019) separated: AI-before became faster than both AI-after and Mixed (p = .02 for both) and carried higher overreliance than both (p = .04 for both). Overreliance on AI-before rose significantly from 0.41 to 0.59 when time pressure was introduced (p = .008), confirming H1 and H1.1. AI-before is therefore the better choice under time pressure on the tradeoff, buying speed at no accuracy cost, but it buys that speed with overreliance.

Overreliance behaves like a stable individual trait. In both experiments, whether a participant overrelied in the first half of the study predicted whether they overrelied in the second (experiment 1: chi-squared(1, N = 93) = 16.28, p < .0001; experiment 2: chi-squared(1, N = 130) = 19.6, p < .0001 without time pressure and chi-squared(1, N = 136) = 16.1, p < .0001 with it). Overreliance was also negatively correlated with response time in both time conditions (p = .0008 and p < .0001), so overreliers are simply faster. H2 and H2.1 were supported. Personality traits were much weaker: no significant predictors, only marginal negative correlations with neuroticism (r = -0.28, p = .07) and Need for Cognition (r = -0.31, p = .10) without pressure, and a marginal positive correlation with self-reported time-pressure performance under pressure (r = 0.35, p = .051). RQ1 returns essentially a null.