← Back to the framework

Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision Making

N = 402 House price estimation

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes Low
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 80%-90%
AI design Manipulated
Protocol Human-first (update)
AI stance Manipulated (Directive, Reflective)
Interactivity Manipulated (Static, Interactive/Dynamic)
XAI Manipulated
Presence Manipulated
Type None · Feature importance · Contrastive · Conversational
Quality Correct
Number of AI advisors One
Research questions
  • RQ1: How does AACT affect people's decision performance and their reliance on AI assistance?
  • RQ2: How does AACT facilitate people's learning of the decision domain?
  • RQ3: How does AACT shape people's experience, perceived critical thinking ability and perceptions of AI in decision making?
  • RQ4: How do individuals with different characteristics respond to AACT differently?

AACT is the authors' own framework, AI-Assisted Critical Thinking, which they introduce in this paper. It inverts the usual direction of analysis: rather than the human evaluating the AI's reasoning, a domain-specific model performs counterfactual analysis of the human's stated decision argument in order to surface flaws in it. The motivating premise is that human-AI team performance remains suboptimal "partially due to insufficient examination of humans' own reasoning".

In the authors' words

"we find that AACT outperforms traditional AI-based decision-support in reducing over-reliance on AI, though also triggering higher cognitive load."

abstract

"our results indicate that AACT is highly effective in reducing decision-makers' over-reliance on AI, but it may also come at a cost of increasing their under-reliance on AI."

results, RQ1

"This implies that only individuals who are more familiar with the task background may benefit from AACT to reduce their over-reliance on AI."

results, RQ4

"we note that these differences are mainly caused by that participants who reported to be very familiar with AI tend to have worse decision performance than other subgroups when receiving no AI assistance or direct decision recommendations from AI - the latter may result directly from their over-reliance on AI."

results, RQ4
Experimental design

A randomised between-subjects online experiment with four treatments, run on Prolific and approved by the authors' IRB.

Task. House sale price prediction from the Ames Housing dataset, pre-processed to 8 features: number of bedrooms, central AC, fireplaces, overall material and finish, kitchen quality, overall condition, age when sold, and living area. Price was discretised into three classes, Low below $100,000, Medium $100,000 to $200,000 and High above $200,000, giving an imbalanced set of 237, 1,837 and 857 instances across 2,930 total. Each participant completed 20 tasks. The authors chose this domain because laypeople can reason meaningfully about it without specialist expertise while still benefiting from assistance, and because the Analyzer treatment requires a multiclass problem.

AI model. Logistic regression on an 80:20 split, reaching 0.874 accuracy on the test set.

Treatments. Human-only: no AI assistance at any point. Recommender: an explicit decision recommendation with LIME feature-importance explanations and the model's confidence, described by the authors as the most common form of decision support in practice. Analyzer: for each of the three possible decisions, LIME-derived evidence for and against it, with no prediction and no confidence shown; this implements Miller's hypothesis-driven XAI, which presents multiple hypotheses simultaneously but does not engage the decision-maker's own reasoning. AACT: a guided conversation in which the AI critiques and helps correct the participant's own stated decision argument, again with no AI prediction, explanation or confidence shown.

What participants do on every task. Provide a decision, select a subset of the task's features as the evidence or argument behind it, and rate confidence. In the three AI treatments they may then update their decision.

Staged procedure. The three AI treatments run in four phases: pre-test of 5 tasks with no assistance, an AI assistance tutorial, 10 intervention tasks with assistance, and a post-test of 5 tasks again with no assistance. The Human-only treatment completes all 20 in one unassisted block. The pre-test and post-test bracket exists to measure learning independently of assisted performance.

Measures. Decision performance as accuracy on tasks 6 to 15, complemented by balanced accuracy given the class imbalance. Reliance as agreement fraction and switch fraction, plus over-reliance, the frequency of following an incorrect AI decision, and under-reliance, the frequency of failing to follow a correct one; for the Analyzer and AACT treatments these were computed against the withheld AI prediction that participants never saw. Learning as normalised change in accuracy from pre-test to intervention (during) and from pre-test to post-test (after). Subjective measures on 5-point scales: NASA-TLX components mental demand, effort and performance, decision confidence, four self-reported critical-thinking abilities (comprehensive evidence evaluation, multiple perspectives, seeking counter-evidence, explaining one's own decision), and AI perceptions covering helpfulness, trustworthiness, understandability, reflection-provocation, decision autonomy, satisfaction and willingness to use. Analyzer and AACT participants were additionally asked whether they would prefer the Recommender instead, and why.

Full findings

AACT bought a large reduction in over-reliance and paid for it in under-reliance and cognitive load, with no net gain in accuracy. The headline is that critiquing the human's own argument moves reliance without moving performance, and the trade is close to symmetric.

Reliance. Participants in the AACT treatment agreed with the AI far less than those given a recommendation (M = 0.684, SD = 0.177 against M = 0.772, SD = 0.188; p = 0.005, d = -0.485), and when their initial decision differed from the model's they switched to it far less often than in either comparison treatment (M = 0.243 against 0.409 in Recommender, p = 0.005, d = -0.461; against 0.393 in Analyzer, p = 0.012, d = -0.423). Decomposed by correctness, over-reliance fell from 0.67 (SD 0.353) in Recommender to 0.543 (SD 0.342) in AACT (p = 0.035, d = -0.365), while under-reliance rose from 0.202 (SD 0.195) to 0.281 (SD 0.191) (p = 0.019, d = 0.409). The authors state the trade-off plainly rather than presenting the over-reliance reduction alone.

Accuracy. There were significant differences across treatments in raw accuracy (F(3,398) = 5.704, p = 0.001, eta-squared 0.041), but they came from Recommender and Analyzer beating the Human-only baseline; AACT did not differ significantly from any other treatment. On balanced accuracy, which corrects for the class imbalance, nothing differed at all (F(3,398) = 1.643, p = 0.179). The authors read this as showing that every form of AI assistance, AACT included, improved overall performance without improving performance on minority-label cases.

Learning. No significant results for RQ2. Neither AACT nor any other treatment improved unassisted performance during or after the intervention, despite the pre-test and post-test bracket built specifically to detect it. This is a clean null on the transfer claim often made for cognitive-engagement interventions.

Cost. AACT participants reported significantly higher mental demand than Human-only (M = 3.887, SD 1.054 against 3.286, SD 1.072; p < 0.001, d = 0.566) and than Recommender (3.33, SD 1.205; p = 0.002, d = 0.493), and higher effort than all three other treatments, with the largest gap against Human-only (d = 0.755). On perceived critical thinking, AACT participants reported no better ability to evaluate evidence comprehensively, consider multiple perspectives or explain their decisions, but did report seeking counter-evidence more than Human-only participants (p = 0.01, d = 0.417) - which is the one thing the system explicitly does for them. Perceptions of the AI itself did not differ across treatments.

The subgroup results are where the paper becomes interesting, and they qualify the main finding heavily. The reduction in over-reliance did not occur in the general population so much as in particular subgroups. Among the task-unfamiliar (n = 195) there was no difference between treatments at all; among the task-familiar (n = 207) the effect was strong (F(2,152) = 8.292, p < 0.001; AACT against Recommender p = 0.001, d = -0.774), and within AACT the task-familiar over-relied less than the task-unfamiliar (p = 0.008, d = -0.537). The same pattern held for AI familiarity: only among those reporting the highest familiarity (n = 93) did over-reliance differ across treatments (F(2,63) = 8.607, p < 0.001, eta-squared 0.215; AACT against Recommender p < 0.001, d = -1.35, the largest effect in the paper), and only participants with a bachelor's degree or above (n = 240) reduced over-reliance through AACT.

For the AI-very-familiar subgroup, AACT also improved performance, reaching significantly higher balanced accuracy than deciding alone (p = 0.037, d = 0.819). The authors are careful about the reason, and it is not flattering to that subgroup: people who reported high AI familiarity performed worse than others both unassisted (p = 0.008) and when given direct recommendations (balanced accuracy p = 0.025), the latter plausibly because they over-relied. AACT helps them most because they had the most room to improve. Consistently, these participants also rated AACT as significantly more reflection-encouraging than Recommender (p = 0.009, d = 0.819).