Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making
N = 258 Trip planningThis study in the framework
Held constant at 66.7% across all six experimental conditions. The level was chosen deliberately: high enough that relying on the system is helpful, low enough that reliance still carries risk, so that the design demands appropriate reliance rather than either blind acceptance or blanket rejection.
- RQ1: How does task complexity influence user trust and reliance on an AI system?
- RQ2: How does task uncertainty, characterized by prognostic versus diagnostic tasks, influence user trust and reliance on an AI system?
- RQ3: How does task complexity interact with task uncertainty to shape user trust and reliance on an AI system?
"individuals under-rely on the AI system in tasks with relatively lower complexity, and over-rely on the AI system in tasks with relatively higher complexity without being able to recognize when the advice may be inaccurate"
results, H1a
"participants' trust in the AI system remains relatively stable regardless of the level of uncertainty in the task"
results, H2b
"when faced with highly complex and prognostic tasks, participants are more likely to relinquish some cognitive control and rely heavily on the AI system"
results, H3
Experimental design
A 3 by 2 between-subjects factorial experiment on Prolific: task complexity (low, medium, high) by task uncertainty (diagnostic, prognostic), giving six conditions referred to as LowDiag, LowProg, MedDiag, MedProg, HighDiag and HighProg. N = 258.
Task. Trip planning. Participants selected the route that minimised both travel time and expense, working from an interface with five components: task scenario and description, a map, route information, general information, and a two-stage decision. Twenty-four scenarios were built, four per condition. Each participant completed three task instances matched to their assigned complexity and uncertainty levels. The domain was chosen because it is a familiar real-world problem and because complexity and uncertainty can both be manipulated within it.
Operationalising complexity. The number of constraints presented, following Wood's component complexity. Low complexity gave four features, medium eight, high twelve. The cut points are grounded in Miller's finding that working memory handles roughly seven plus or minus two chunks: five to nine features count as medium, more than nine as high, four or fewer as low.
Operationalising uncertainty. Diagnostic tasks placed the trip in the present and gave precise values for every constraint, eliminating ambiguity. Prognostic tasks placed the trip two weeks in the future and replaced exact values with ranges or estimates, adding probabilities for certain outcomes such as a high likelihood of rush-hour congestion or a low chance of rain. Task features were designed to be independent of one another and were classified as time-dependent (traffic, weather) or time-independent; low-complexity tasks balanced the two, while medium and high complexity increased the proportion of time-dependent features.
AI system. Tuned to 66.7% accuracy in every condition, chosen so that relying on it is helpful but still carries risk, requiring appropriate reliance rather than blind acceptance. Within each participant's batch of three tasks, incorrect advice was given exactly once, at a random position, controlling for order effects.
Measures. Performance as accuracy. Reliance as agreement fraction and switch fraction. Appropriate reliance as accuracy on trials of initial disagreement (Accuracy-wid), relative positive AI reliance (RAIR) and relative positive self-reliance (RSR). Trust via the Trust in Automation questionnaire, covering reliability and competence, understanding and predictability, intention of developers, and overall trust, on 5-point scales. Covariates were the Subjective Numeracy Scale, the Affinity for Technology Interaction scale, TiA-Familiarity and TiA-Propensity to Trust.
Full findings
Complexity and uncertainty both moved reliance, and both moved it in the wrong direction: participants leaned harder on the AI exactly where they were least able to tell good advice from bad.
Complexity first. Switch fraction rose monotonically across the three levels (adjusted p = .003; low 0.18, medium 0.26, high 0.34, with low < medium < high), while agreement fraction did not differ (p = .8). Accuracy fell (p < .001; 0.79, 0.58, 0.61, low above both others). Decomposing appropriate reliance, Accuracy-wid fell (p = .001; 0.61, 0.40, 0.50) and RSR fell sharply (p < .001; 0.64, 0.34, 0.38) while RAIR rose (p = .001; 0.22, 0.33, 0.43). The authors read the rise in RAIR carefully and correctly: it does not mean reliance became more appropriate. It means participants under-relied on the AI when tasks were simple and over-relied when they were complex, without gaining any ability to spot bad advice. H1a was partially supported.
Uncertainty produced the same pattern on every measure. Prognostic tasks lowered agreement fraction (p = .01; 0.60 against 0.54) but raised switch fraction (p = .02; 0.22 against 0.31), lowered accuracy (p < .001; 0.72 against 0.60), lowered Accuracy-wid (p = .04; 0.56 against 0.45) and raised RAIR (p = .02; 0.27 against 0.38). RSR did not differ significantly (p = .1). H2a was partially supported. The authors take the switch-fraction result as evidence that people can read the uncertainty of a task and adjust reliance accordingly; the Accuracy-wid result shows that the adjustment does not help them.