← Back to the framework

Don't Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI Explanations

N = 204 Logical / reasoning task

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty Not reported
AI Ecosystem Designable
AI performance 95%-100%
AI design Manipulated
Protocol AI-first
AI stance Manipulated (Directive, Reflective)
Interactivity Static
XAI Manipulated
Presence Manipulated
Type None · Narrative
Quality Correct
Number of AI advisors One
Research questions
  • RQ1: Do humans perform better at discerning the logical validity of socially divisive statements when they receive feedback from AI systems compared to when they work alone?
  • RQ2: How do AI-framed Questioning and causal AI-explanations affect participants' discernment of logical validity, confidence of their discernment, perceived information sufficiency by controlling personal factors (i.e. prior belief, trust in AI, cognitive reflection) as covariates?
  • RQ3: Do personal factors, such as prior belief, prior knowledge, trust in AI, cognitive reflection (indicating the level of critical thinking) impact discernment?

The study was pre-registered at aspredicted.org under #94860 before being conducted.

In the authors' words

"Our results show that compared to no feedback and even causal AI explanations of an always correct system, AI-framed Questioning significantly increase human discernment of logically flawed statements."

abstract

"Since AI-framed Questioning is prediction agnostic and always gives the same question feedback - whether or not a statement is logically valid or not - the user is unable to rely on the prediction of the AI system (there simply are no answers given)."

discussion, limitations

"Such a finding suggests that individuals tend to find the given information is sufficient enough to support the claim (as measured by a significantly lower perceived information insufficiency) when their judgement is corroborated by a second opinion from AI in the causal explanation form."

results, 5.3
Experimental design

A 3-by-2 factorial design: three feedback conditions between-subjects, crossed with the logical validity of the statement within-subjects.

Materials. Statements were drawn from the IBM Debater Claims and Evidence dataset, which pairs a claim with supporting evidence across 58 socially divisive topics. Five topics were sampled at random: violent video games cause aggression; affirmative action counters the effects of a history of discrimination; refugees should be embraced; Israel should lift the blockade of Gaza; male infant circumcision should be less prevalent. Within each topic five anecdotal and five non-anecdotal claim-plus-evidence pairs were sampled. Anecdotal support commits the hasty generalisation fallacy, so those pairs were labelled logically invalid and the non-anecdotal ones logically valid. The authors then corrected every statement against a formal definition of validity, ending with four valid and four invalid statements per topic, 40 in total. Linguistic markers that would leak validity were stripped out, for instance names and phrases such as "researchers show" or "most studies", and the resulting sets did not differ significantly on word count, Flesch-Kincaid grade level or sentiment.

Task. Each participant judged 10 statements sampled at random from the 40. After reading a statement they clicked next, the AI feedback slid up, and they then reported whether the statement was logically valid, how confident they were on a 1-7 scale, and whether sufficient information was given in the statement to support the claim, also 1-7.

Conditions. Causal AI-Explanation: the system states a reason for the label, "If X then it follows that Y" for valid statements and "If X then it does not follow that Y" for invalid ones. AI-framed Questioning: the system asks about the same causal link without indicating whether the label follows, "If X does it follow that Y?", identically phrased for valid and invalid statements. No-Explanation: no feedback of any kind. Explanations in both AI arms were generated with GPT-3 from a small set of hand-crafted templates and then manually checked for accuracy and consistency, and checked to ensure no linguistic difference between the two arms beyond the statement-specific reason and label.

Measures. Weighted discernment of logical validity, a 0-100 continuous score combining binary accuracy with the confidence rating so that a confidence of 1 pulls the judgement to the neutral midpoint of 0.5 and a confidence of 7 leaves it at 0 or 1. Perceived information insufficiency, the inverted sufficiency rating, where 7 means the participant found the information insufficient and would seek more. Covariates measured were prior belief and prior knowledge per topic, trust in AI through the six-item Mayer, Davis and Schoorman ability-benevolence-integrity battery as used by Epstein et al., and cognitive reflection through three items sampled from the extended CRT. Completion time was logged. Free-text reports of participants' thinking were collected and inductively coded by two coders to saturation.

Full findings

Asking a question beat giving the answer. On the statements that mattered, the logically invalid ones, participants who were asked a Socratic question discerned validity better than participants given a correct causal explanation, and both beat working alone.

Unaided performance was the baseline problem: on invalid statements the control group reached 44% raw accuracy (SD = 26), below the 50% expected from guessing, meaning participants were not merely uninformed but systematically taken in by anecdotal evidence. Raw accuracy rose to 57% with causal explanations and 67% with AI-framed Questioning.

On weighted discernment for invalid statements the MANCOVA gave a significant main effect of condition, F(2, 1007) = 15.3, p < .001. Questioning scored 62.5 (SE = 1.9), causal explanation 55.0 (SE = 2.2) and control 46.9 (SE = 2.1); questioning beat control at p < .001, causal explanation beat control at p = .007, and questioning beat causal explanation at p = .009, all against Benjamini-Hochberg-adjusted thresholds. H2 is supported where it counts.

On valid statements the ordering changes and the questioning advantage disappears: condition again mattered, F(2, 1007) = 8.4, p < .001, but causal explanation (78.2, SE = 1.7) and questioning (74.7, SE = 1.6) both beat control (68.2, SE = 1.8) without differing from each other, p = .126. This asymmetry is the substantive result. A correct causal explanation is at least as good as a question when it confirms a sound argument; it is worse when the argument is fallacious, which is exactly the case where independent scrutiny is needed.

The mechanism the authors propose is that questions have no truth value: because the questioning system is prediction-agnostic and issues the same question whether the statement is valid or invalid, there is no answer to defer to, so the participant must resolve it themselves. The complementary evidence is in perceived information insufficiency. Only the causal explanation arm lowered it, to 4.4 (SE = 0.1) against control 4.9 (SE = 0.1), p = .002 on invalid statements and p < .001 on valid ones, and it also sat significantly below the questioning arm on valid statements, p < .001. Being told why, by a system that is always right, makes people feel the evidence in front of them is sufficient and stop looking. Being asked does not.