The Impact of Placebic Explanations on Trust in Intelligent Systems
N = 30 Nutrition recommendationThis study in the framework
Not applicable. There is no model and therefore no accuracy figure. The prototype was a click-dummy mobile app in which the personalised recommendations were simulated rather than computed, chosen deliberately so that the three arms would be comparable and would see identical outputs. Explanations were authored by the researchers, not generated. The recommendation set was fixed: one recommended meal plus two alternatives at each of breakfast, lunch and dinner, with the lunch recommendation always a salad. There is no ground truth against which any of these could be scored correct or incorrect, so neither the AI's accuracy nor the participants' decision accuracy is defined.
RQ: Do placebic explanations invoke a similar level of trust in an intelligent system as real explanations?
"Our results indicate that placebic explanations for algorithmic decision-making may indeed invoke perceived levels of trust similar to real explanations."
abstract
"it is interesting to see that both the understandability as well as the scope and content of the explanations were perceived as being similarly sufficient in both the placebic- and real-explanation group."
results
"This placebo effect seems potentially worrisome if used to "deceive" users in a sense comparable to "dark" UX patterns."
discussion
"future work on explanations for intelligent systems might consider using a placebic explanation as a baseline, not (only) a baseline without any explanations at all."
discussion
Experimental design
Between-groups lab study with three arms, one per explanation type: no explanation, placebic explanation, real explanation. N = 30, randomly assigned, 10 per arm.
Task. Participants were given a written scenario asking them to build a three-meal nutrition plan for their mother, a 47-year-old nurse who wants to lose 4 to 6 kg over three months and who "likes salad the least". Deciding on behalf of someone else was deliberate, to remove participants' own food preferences from the decision. Participants completed an onboarding flow entering the mother's age, weight, height and dietary goal, then received one recommended meal plus two alternatives for each of breakfast, lunch and dinner and chose one of the three. For lunch the prototype always recommended a salad, so that the recommendation conflicted by construction with the stated preference and the choice functioned as a behavioural test of compliance.
AI system. A click-dummy mobile app in German, with no implemented model. Recommendations were simulated so that all three arms saw identical outputs. Explanations were authored rather than generated.
Manipulation. The placebic explanations introduce a justification with because, since or so that, but add nothing beyond what the scenario already implies, for example "We need these details because they are necessary for the algorithm". The real explanations describe the apparent computation and, at the recommendation step, state the system's certainty, for example that the number of calories and nutritional values "has been calculated based on your personal details". Explanations appeared in both the onboarding part and the recommendation part. The no-explanation arm saw the same prototype with no explanatory text at all.
Measures. A post-task questionnaire on 5-point Likert scales from 1, strongly disagree, to 5, strongly agree, taken from Corritore et al. and extended by the authors with three items on perceived understanding of the algorithmic decision-making. Meal choice was recorded, in particular whether the participant took the recommended salad at lunch. Participants were then asked orally for the reasons behind their choices.
Full findings
Placebic explanations produced trust ratings indistinguishable from real explanations, and both sat above the no-explanation arm. This is the paper's central claim, and one the authors are careful to present as a tendency rather than a demonstrated effect.
Across the trust items, group medians in the placebic and real arms exceeded the no-explanation arm on every question but one, and the two explanation arms did not separate from each other. The same held for the two meta-items: participants rated the explanations similarly understandable, and rated their scope and content similarly sufficient, whether or not the explanation carried any information. That second result is the sharper one, because it says the vacuity of a placebic explanation is not detected by the people receiving it.
Meal choice moved in the same direction and produced a cleaner contrast than the ratings. None of the 10 participants in the no-explanation arm chose the recommended salad, and all 10 cited the mother's dislike as the reason. Four of 10 chose it in the placebic arm and four of 10 in the real arm. The stated reasons diverged, however. In the placebic arm those who complied said they wanted the best possible result for the mother and therefore weighted the algorithm's recommendation above her preference; in the real arm they pointed specifically to the high stated system certainty. Among those who declined, three placebic-arm participants said the information was insufficient to override the preference, while five real-arm participants said the advantage over the alternatives did not look large enough to be worth it. The same behavioural rate was therefore reached by two different routes: a contentless justification moved four participants without giving them anything to reason with, and a real explanation moved four by giving them a figure to weigh.
The authors draw the methodological consequence explicitly. An experiment that contrasts explanation against no explanation cannot separate the effect of explanation content from the psychological effect of being given a justification at all, so a placebic arm should serve as the baseline. They also flag the deployment risk, comparing an empty explanation offered to soothe users to a dark UX pattern.