Effects of Interaction and Explanation Type on Human–AI Collaboration for Fake News Detection
N = 161 Deception detectionThis study in the framework
Held constant at 75%, and the two figures agree for once. The XGBoost classifier reaches 75% accuracy on the FakeNewsAMT test set, and the 13 articles drawn from that test set for the study were selected so that the assistant was right on 6 of the 8 Categorization Task items, giving participants an experienced accuracy of 75% as well. Selection was constrained on readability (300 to 600 words), topical diversity and classification accuracy, so the item set is curated rather than random, but it is curated to reproduce the model's own rate rather than to inflate it.
- RQ1. How does the interactivity of the AI assistant affect the decision-making process and human–AI collaboration?
- RQ2. How do different explanation types influence users' decision performance and their reliance patterns in the AI system?
- RQ3. What are the carry over effects of AI system design on users' decision-making when they later perform the task without AI assistance?
No numbered hypotheses are stated. All three questions are answered with human participants; there is no simulation component beyond the model-selection benchmark reported in the system section.
"Results show that conversational systems increased user reliance, while feature importance explanations appeared to be more effective than counterfactual explanations."
abstract
"Although conversational systems are often positioned as tools to promote deeper user engagement and reflection, our study suggests that they may primarily serve a persuasive function, encouraging agreement with the AI, rather than enhancing understanding or improving outcomes."
discussion, RQ1
"The results suggest that users may not intuitively grasp the contrastive nature of counterfactual explanations, which are designed to answer implicit 'what if not?' questions by showing how minimal changes could alter the decision."
results, user engagement with the conversational system
"Notably, neither explanation type nor interactivity level of the AI system showed significant effects on the self-reported perception measures. In contrast, the covariates, participants' predefined AI Profile and News Consumption Profile, were more strongly associated with users' trust and confidence."
results, self-reported perceptions
Experimental design
A 2 × 2 between-subjects design crossing explanation type with interactivity, followed within participants by an unaided phase used to test carry-over.
Conditions. Explanation type was either feature importance, a ranked textual list of the features that most influenced the classification, or counterfactual, a set of feature-based rules describing what would have to change for the prediction to flip. Interactivity was either a static system, where the suggestion and explanation appeared in a fixed non-interactive format, or a conversational system, where the same explanation was delivered through a chatbot that answered follow-up questions. There is no no-explanation control arm, which the authors state as a deliberate choice: the benefit of explanations over none is taken as established, and the study was designed to compare explanation types and modalities rather than revisit that question.
Task. Binary fake news detection over short fabricated and legitimate news articles of 300 to 600 words. The study ran in two phases. In the Categorization Task participants classified eight articles with the assistant, called NewsBot, present. In the Evaluation Task they classified five further articles with no assistance and wrote a free-text justification for each. Manipulations applied only to the Categorization Task; the Evaluation Task is the carry-over measure. In both phases the response options were Fake News, Not Fake News and I Don't Know. The third option was included in place of a binary forced choice because the design does not record a pre-advice judgement: with the reward structure attached, choosing I Don't Know marks hesitant non-reliance, neither endorsing nor overriding the system. Participants were barred from browsing or using phones and were timed out after two minutes of inactivity on an article.
AI system. An XGBoost classifier over roughly 20 engineered psycholinguistic features selected from more than 100 indicators covering linguistic complexity, emotional valence, cognitive markers, social references and contextual anchors, extracted with LIWC, MoralStrength, TextBlob and VADER. Trained on FakeNewsAMT, 480 articles balanced between fake and legitimate. XGBoost was chosen over CatBoost, LightGBM, Random Forest, logistic regression and SVM on the same split; the tree ensembles reached 71 to 72% and the linear models 57 to 61%.
Explanations. Feature importance explanations come from TreeSHAP, truncated to the top five attributions because those account on average for more than 70% of total attribution magnitude, and presented as a ranked list showing direction but not numerical values. Counterfactual explanations come from LORE, whose local decision-tree surrogate yields a factual rule and a set of minimally perturbed counter-rules; only the direction of the required change is rendered into natural language, again with values omitted. Both omissions are stated as cognitive-load control. The conversational assistant is GPT-4 at temperature 0, constrained by defensive prompting to paraphrase the pre-computed TreeSHAP or LORE explanation without introducing new reasoning or content, with a canonical refusal for off-topic queries and a fresh context per article.
Incentive. £4 base plus a performance bonus of £0.08 per correct classification and £0.04 deducted per incorrect one, floored at £0, with I Don't Know neither rewarded nor penalised; approximately £10.50 per hour overall. The authors state the structure exists to simulate stakes and to make the three-option response interpretable as reliance behaviour.
Measures. Objective, computed separately for each phase: classification accuracy, average decision time, agreement rate with the AI, disagreement rate excluding I Don't Know responses, and I Don't Know rate. Self-reported: a three-item human confidence composite covering confidence in one's own classifications in each phase and in NewsBot's recommendations (Cronbach's alpha = .73), and the Multidimensional Measure of Trust across reliability, competence, ethics, transparency and benevolence (alphas .71 to .94). Covariates: an AI Profile composite of four items on AI use, familiarity, comfort and general trust (alpha = .79), and a News Consumption Profile composite of four items on reading frequency, sharing, and trust in mainstream media and in social contacts (alpha = .69).
Full findings
The conversational interface made participants agree with the AI more without making them any more accurate, which is the paper's central result and its clearest warning.
Interactivity produced significant main effects on agreement rate and on I Don't Know rate, both surviving Holm-Bonferroni correction: participants using the conversational system agreed with the recommendation more often and declined to commit less often than those seeing the same explanation statically. Accuracy did not follow. The authors read the combination as uncalibrated reliance and are direct about the implication for a technology usually justified on the opposite grounds: conversational systems may primarily serve a persuasive function, encouraging agreement with the AI, rather than enhancing understanding or improving outcomes. Decision time was also longer in the conversational conditions, which the authors treat as expected from the interaction format, but that effect did not survive correction.
Explanation type separated the two families in the direction of the simpler one. Counterfactual explanations produced lower accuracy, longer decision times and higher disagreement rates than feature importance explanations; after correction only the decision time difference held. The chatlog analysis supports the reading that counterfactuals were harder to use rather than merely less liked: participants in the counterfactual conditions asked significantly more clarification questions, particularly why questions, which the authors take as evidence that users do not intuitively grasp the contrastive what-if-not structure counterfactuals are built to answer. No interaction effects between explanation type and interactivity were found on any measure.
The carry-over result is narrow and behavioural. In the unaided Evaluation Task, participants previously exposed to the conversational system selected I Don't Know less often, a significant main effect of interactivity. Nothing else carried over: explanation type had no effect on any unaided measure, and unaided accuracy did not differ by condition. Since there is no pre-exposure baseline, the design shows a difference between conditions in post-exposure unaided behaviour, not an improvement attributable to exposure.
The most consequential null is on the subjective side. Neither explanation type nor interactivity affected any self-reported measure, not one MDMT trust subscale, not the confidence composite. Both covariates, in contrast, were significantly associated with every self-reported variable: participants with more AI experience and heavier, more trusting news consumption reported higher trust across all five trust dimensions and higher confidence in themselves and in the system. The same covariates were associated with none of the objective measures in either phase. Stated trust in this study is a property the participant brought with them, not one the interface produced, while behaviour moved with the interface and not with the participant. That dissociation is the finding most likely to transfer, and it argues against using self-reported trust as a proxy for reliance in this task class.