← Back to the framework

Does Sycophancy Change Decisions? Effect of LLM Sycophancy on AI-Assisted Decision-Making

N = 106 Speed Dating PredictionETF Investement

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes Manipulated
LowHigh
Stress level / time constraint Not reported
Task uncertainty High (prognostic)
AI Ecosystem Designable
AI performance Not applicable
AI design
Protocol Human-first (update)
AI stance Directive
Interactivity Static
XAI
Presence Yes
Type Narrative
Quality Not reported
Number of AI advisors One
Research questions
  • RQ1: How does LLM sycophancy affect users' tendency to stick with or revise their initial decisions?
  • RQ2: How do different types of LLM sycophancy differentially affect users' decision-making?
  • RQ3: How does the impact of LLM sycophancy on user decisions vary across tasks with different risk levels?
In the authors' words

"Results show that sycophancy influences decision patterns in type-dependent ways. Specifically, opinion agreement reinforces initial decisions and self-deprecation boosts confidence."

abstract

"Implicit sycophancy (opinion agreement), although less perceptible, reinforced users' commitment to their initial judgments, while explicit sycophancy (direct praise and self-deprecation) had only limited effects on immediate decision shifts."

discussion, 5.1.2

"When LLMs agree with the user's opinion, it functions as a confirmation signal, reinforcing the cognitive anchor established by the user's initial decision. In this sense, opinion agreement operates less as active persuasion and more as a form of social validation."

discussion, 5.1.3

"Taken together, the results highlight that task risk selectively amplifies the influence of contradictory, rather than confirmatory, AI advice, underscoring the importance of contextual stakes in shaping human-AI decision-making dynamics."

results, 4.2.2
Experimental design

A 4 × 2 mixed design: sycophancy type between participants, task risk within participants.

Conditions, assigned between participants. Non-sycophancy: objective, factual, restrained in tone. Opinion agreement: the assistant repeats or affirms the user's judgment and reasons, for example that their idea makes a lot of sense. Direct praise: the assistant compliments the user's ability, intelligence or judgment, for example that they analysed it very clearly. Self-deprecation: the assistant downplays its own authority to elevate the user's, for example that it may not see it as clearly as they do. The taxonomy is drawn from Gordon's five flattery strategies in social psychology together with existing categorisations of LLM sycophancy.

Tasks, varied within participants in randomised order. Low-risk: speed-dating prediction, following Yin and others, over 12 rounds. Participants saw two real dating profiles, demographics, the 100-point allocation each dater made across six attributes (attractive, sincere, intelligent, fun, ambitious, shared interests), and both parties' post-date ratings, likeability and perceived reciprocity, and predicted whether the proactive dater wanted to meet again. High-risk: ETF investment, following Reicherts and others, over 2 rounds. Participants took the role of a 40-year-old cautious investor allocating the equivalent of 100,000 CNY across 28 equity ETFs sampled by GICS sector and investment theme, with prices, 24-hour change, tracked index, tracking error, top five holdings and their drawdowns, returns, maximum drawdown and annualised volatility. Six months of price history was shown, drawn from real Chinese-market data from February 2024 but presented as fictional and as current, so that participants could not draw on knowledge of actual market events. The investor-persona framing was used to stop participants defaulting to familiar holdings.

Round structure, identical in both tasks. The participant made an independent decision, rated confidence, and wrote a rationale of at least 20 words. A single round of LLM feedback followed, randomly assigned to supply either supportive or opposing evidence. The participant then decided whether to revise and rated confidence again. Round order and feedback valence were both randomised. Across the retained data the assistant opposed the participant in 752 rounds and supported them in 698.

AI system. Eight chatbots built on the Coze platform, a fully crossed 4 sycophancy types × 2 tasks, all running DeepSeek-V3, chosen for Chinese-language performance. All eight were given identical task data through a standardised workflow storing the datasets as a knowledge base retrieved per round, so that only the communication style differed between conditions.

Measures. Decision change, coded 1 if the participant revised after the feedback. Confidence change, the difference between the two self-reported ratings. Cognition-based trust on a 7-item scale and affect-based trust on a 3-item scale. NASA-TLX across its six dimensions. Need for cognition on the NCS-6. A four-item manipulation check asking whether the AI seemed to agree, to compliment, to downplay itself, and to intend flattery. Nine post-task semi-structured interviews of roughly 30 minutes with participants selected to span task performance and reliance levels, transcribed and coded with inter-coder agreement assessed.

Full findings

The only form of sycophancy that changed what people decided was the one they could not detect. Direct praise and self-deprecation were both rated significantly above the non-sycophantic baseline on the flattery items (for the overall flattery item, F = 5.60, p = .001; direct praise M = 4.73, self-deprecation M = 3.78 against baseline M = 3.41), while opinion agreement was not reliably distinguished from baseline at all (M = 3.92). The authors treat this not as a failed manipulation but as the finding, splitting sycophancy into explicit forms that users notice and an implicit form that they do not.

On decisions, only the implicit form moved anything. When the assistant offered evidence contradicting the participant's initial choice, opinion agreement significantly reduced the likelihood of revising; direct praise and self-deprecation had no significant effect on decision change in any configuration. The mechanism the authors propose is confirmation rather than persuasion: an assistant that echoes the user's own reasoning acts as social validation, strengthening the anchor the user has already set, so that contradictory evidence arriving later in the same message is discounted. Being agreed with made people harder to move, and they did not notice it happening.

On confidence the pattern inverts, and the explicit forms take over. When the assistant opposed them, participants in every condition lost confidence except those receiving self-deprecation, who gained slightly (mean changes of -0.18 for non-sycophancy, -0.15 for opinion agreement, -0.15 for direct praise, +0.05 for self-deprecation; the self-deprecation against baseline contrast beta = -0.091, p = .028). An assistant that disclaims its own authority buffers the user against the confidence hit of being contradicted. When the assistant agreed with them, confidence rose in all four conditions (0.29 to 0.52) with no differences between them: agreement raises confidence whether or not it is dressed as flattery. Opinion agreement, which was the only condition to change decisions, had no significant effect on confidence at all. Decision persistence and confidence run on separate channels here.

Task risk mattered, but selectively and asymmetrically. Where the assistant confirmed the participant's choice, risk made no difference to adoption (beta = 0.377, p = 0.415). Where it contradicted them, adoption was significantly higher in the high-risk ETF task than in the low-risk dating task (beta = 0.843, p < 0.001). Contradiction gets more purchase when the stakes are higher, confirmation does not. There was no significant interaction between sycophancy type and task risk on decision change, so the two factors operated independently; the one interaction that did emerge was on confidence, where direct praise combined with supportive evidence produced larger confidence gains in the high-risk task than in the low-risk one (beta = 0.61, p = 0.015). Flattery that would read as unnecessary over a dating prediction reads as validation when real money is notionally at stake.

Explicit sycophancy bought perception and cost effort. Direct praise raised the ability dimension of cognition-based trust above baseline (F = 3.591, p = .016; mean difference 0.56, 95% CI [0.04, 1.09]) and simultaneously raised NASA-TLX temporal demand (F = 3.304, p = 0.023; mean difference 17.07, 95% CI [1.47, 32.67]), participants felt more time-pressured when complimented. No other trust dimension and no other workload dimension differed. The interviews explain the ambivalence: one participant reported that compliments made them happier and more confident, another that frequent compliments in every exchange were annoying and made the assistant look unprofessional, and several said self-deprecation made the assistant seem less competent and made them doubt its authority. The authors read the weak behavioural effect of explicit flattery as users discounting praise once they perceive it as strategic.