Decision control and explanations in human-AI collaboration: Improving user perceptions and compliance
N = 483 Hotel price estimationThis study in the framework
Held constant, and never disclosed to participants. A model trained with five-fold cross-validation on hotel data, reaching a mean absolute error of 21.29 pounds on the test set, described as consistently high performance across training folds and random initialisation seeds. Room prices across the six hotels ranged from 144 to 316 pounds, so the error is roughly a tenth of the price. The authors note that given this accuracy it would be in the participants' best interest to accept the recommendation in most cases, but state explicitly: "we do not communicate the model's accuracy". Features such as review texts were removed to avoid amplifying racial bias.
H1a. High decision control improves user perceptions of (1) trust in, and (2) understanding of the system recommendation. H1b. High decision control improves users' (1) intended, and (2) actual compliance with the system recommendation. H2a. Explanation presence improves user perceptions of (1) trust in, and (2) understanding of the system recommendation. H2b. Explanation presence improves users' (1) intended, and (2) actual compliance with the system recommendation. H3a. Explanation presence increases perceived task complexity, which in turn impairs users' (1) trust in, and (2) understanding of the system recommendation. H3b. Explanation presence increases perceived task complexity, which in turn impairs users' (1) intended, and (2) actual compliance with the system recommendation. H4a. The higher the cognitive ability, the stronger explanation presence improves users' (1) trust in, and (2) understanding of the system recommendation. H4b. The higher the cognitive ability, the stronger explanation presence improves users' (1) intended, and (2) actual compliance with the system recommendation.
"users benefit from enhanced decision control, while explanations - unless appropriately designed for the specific user - may even harm user perceptions and compliance"
abstract
"if the cognitive ability was only moderate, providing no explanation improved (rather than some explanation impaired) participants' actual compliance"
results, H4b.2
"the specific user's cognitive ability determines the effectiveness of a provided explanation"
discussion
Experimental design
Three between-subjects online experiments on Prolific, all sharing one task and one platform, analysed separately. Participants were paid a flat $2 for a 15-minute study and all had at least some work experience in the hospitality and tourism sector. Sample sizes after exclusions were 110 (study 1), 110 (study 2) and 263 (study 3). Sample sizes were set a priori with G-Power, assuming a moderate effect of .15 in studies 1 and 2 and a small effect of .10 in study 3.
Task. Hotel revenue management. Participants estimated the nightly room price of six London hotels from 13 hotel characteristics, and were told the average London room price was 180 pounds (Task I, initial decision). They were then shown a machine learning model's predicted price for each hotel and estimated again (Task II, updated decision). The two-stage structure is what makes every study human-first: an independent judgement is recorded before the model is seen.
Study 1 manipulated decision control, low against high. In the low condition participants could only choose between their own initial estimate and the model's recommendation. In the high condition they could adjust their estimate to any value. A separate manipulation check (N = 60, five excluded) confirmed the manipulation: perceived control 4.20 (SD 0.75) in high against 3.64 (SD 1.09) in low, p = .033.
Study 2 held decision control at high and manipulated explanation presence against absence. Explanations were simplified Shapley values communicating the marginal impact of each hotel characteristic on the recommended price. A pretest (N = 76) found that 66% objectively understood the explanation while only 15% felt they fully understood it, and the design was simplified before study 2 on the basis of that feedback.
Study 3 repeated study 2 with three arms: explanation Style A (characteristics in the arbitrary Task I order), explanation Style B (characteristics ordered by size of marginal impact) and explanation absence. Style A against Style B is a robustness check on explanation format; the two were pooled once no difference emerged. Study 3 added cognitive ability, measured before the task with the four-item visual processing scale of the Cattell-Horn-Carroll model, and perceived task complexity, measured after Task II with four items including 'This task was mentally demanding'. Study 3 ran in six blocks with different hotels to prevent order and anchoring effects.
Measures. Trust and perceived understanding, three items each on 7-point scales. Intended compliance as two intention-to-use items. Actual compliance as the relative change in Mean Relative Absolute Error of the participant's estimates towards the model's recommendation from Task I to Task II, benchmarked against always predicting the London average, and sign-reversed so that +1 is maximal compliance. Analysis was multiple regression for each outcome, controlling throughout for gender, age, education and three kinds of work experience, plus a moderated mediation model (Hayes model 5) in study 3.
Full findings
Giving users control over the recommendation improved every outcome; giving them an explanation of it improved none, and the explanation harmed outcomes through the perceived complexity it added.
Decision control (study 1). High decision control raised trust (p = .004, coefficient 0.69) and understanding (p = .004, coefficient 0.71), and raised both intended (p = .029, coefficient 0.52) and actual compliance (p = .010, coefficient 0.17). H1a and H1b were fully supported. A post hoc mediation showed the effect on actual compliance ran through intended compliance (p = .007) while controlling for trust and understanding. The content of the advice was identical across conditions; only the freedom to alter it changed.
Explanation presence (studies 2 and 3). Providing the explanation moved nothing. In study 2, trust p = .500, understanding p = .628, intended compliance p = .431, actual compliance p = .378. Study 3 replicated the null across both explanation formats: trust p = .770 and p = .697, understanding p = .674 and p = .177, actual compliance p = .934 and p = .172, with Style B marginally reducing intended compliance (p = .053). H2a and H2b were not supported anywhere, and the null is robust to the format of the explanation.
The mechanism (study 3). Explanation presence significantly increased perceived task complexity (p < .001, coefficient 0.81), and that increase impaired understanding (p < .001), marginally impaired trust (p = .061), and impaired intended compliance (p = .011). It did not carry through to actual compliance (p = .753). H3a was partly supported, H3b half supported. This is the only measured mediator in the base: the harm is not asserted from a null but traced through a construct the authors instrumented.
Cognitive ability as boundary condition (study 3). The effect of explanations was conditional on the individual. Cognitive ability moderated the explanation effect on understanding (p = .014) and marginally on trust (p = .055): at high cognitive ability the explanation improved understanding (p = .026), at moderate ability it did not (p = .181). For actual compliance the moderation was significant (p = .035) and runs the other way at the bottom of the range: at moderate cognitive ability, withholding the explanation improved actual compliance (p = .044), while at high ability presence and absence converged (p = .334). H4a.2 was supported, H4a.1 partly, H4b.1 and H4b.2 not.