Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating XAI Systems
N = 296 Nutrition estimationThis study in the framework
Simulated and held constant at 75% across every study and condition. There is no underlying model; the recommendations and explanations were authored. Errors were placed at fixed question positions (4, 7, 11, 16, 22 and 23 of 24), so every participant encountered the same number of errors at the same points in the sequence, while which food item carried the error was randomised. This gives identical error incidence across participants with varied error content.
- RQ1: Do results obtained on a proxy task reproduce on the corresponding real decision task?
- RQ2: subjective ratings collected during a real decision task predict objective performance on that same task?
"evaluations with proxy tasks did not predict the results of the evaluations with the actual decision-making tasks"
abstract
"the subjective measures on evaluations with actual decision-making tasks did not predict the objective performance on those same tasks"
abstract
"by employing misleading evaluation methods, our field may be inadvertently slowing its progress toward developing human+AI teams that can reliably perform better than humans or AIs alone"
abstract
"Because [the AI] showed similar pictures, I knew that it had data backing it up"
qualitative study, participant P3
Experimental design
Three studies plus replications of the two online experiments, all built on the same nutrition task and the same two explanation designs.
Task material. Twenty-four images of plates of food. The underlying question is whether X% or more of the nutrients on a plate is fat. A simulated AI answers it, with 75% accuracy: errors were fixed at questions 4, 7, 11, 16, 22 and 23, while which food carried the error was randomised across participants.
The two explanation designs, both within-subjects and counterbalanced, differ in the mode of reasoning they demand rather than in their content domain. Inductive explanations are example-based, opening with a statement that these are plates the AI knows the fat content of and categorises as similar to the one shown, followed by four further food images, so the participant must recognise the contributing ingredients and draw their own conclusion. Deductive explanations state the general rules the AI applied, opening with a statement that these are the ingredients the AI recognised as main nutrients, followed by a list. Crucially, when the AI erred the explanation revealed the error faithfully, showing the wrong similar images or the misrecognised ingredients, for instance churros with chocolate misrecognised as sweet potato fries with BBQ sauce.
Study 1, the proxy task (N = 183 of 200 recruited). Participants were shown the true fat content as a fact and asked what the AI would decide, given the explanations. The design measures whether people build correct mental models of the AI, deliberately removing their own nutrition knowledge from the equation. Within-subjects on explanation type, 24 questions in two blocks of 12. Measures: performance as the percentage of correct predictions of the AI's decision; appropriateness of the examples or ingredients, binary, after every question; trust on a 5-point scale at the end of each block; mental demand on a 5-point scale every four questions; and forced comparisons between the two AIs at the end. MTurk, US adults, mean 7 minutes, $2.
Study 2, the actual decision-making task (N = 102 of 113 recruited). The same 24 images, but participants made their own judgement about fat content, assisted by a simulated AI recommendation. Three between-subjects conditions: a no-AI baseline (n = 18), a no-explanation baseline giving the recommendation alone (n = 19), and the main condition giving recommendation plus explanation (n = 65), within which explanation type varied within-subjects. Measures: performance overall and specifically on questions where the AI erred; perceived understanding after every question; trust every four questions; helpfulness at the end of each block; and forced comparisons. MTurk, US adults, mean 10 minutes, $5.
Study 3, the think-aloud study (N = 11). The same design as the actual decision-making task, main condition only, run in person with screen and audio recording. Participants verbalised their reasoning on each decision, followed by a semi-structured interview about how they believed each AI worked and why they did or did not trust it. Recruited through community mailing lists: 8 women and 3 men, aged 23 to 29, mostly graduate students from design, biomedical engineering and education, with 0 to 5 years of AI experience. Transcripts were coded inductively for how the AI made its recommendations, trust, erroneous recommendations, and reasons for preferring one explanation type.
Full findings
The proxy task and the real task gave opposite answers, and subjective preference on the real task pointed away from actual performance. Taken together the three studies form a double dissociation that undermines two standard evaluation practices at once.
On the proxy task, where participants predicted what the AI would decide, inductive example-based explanations won on every subjective measure. They were trusted more (M = 3.55 against 3.40, F(1,182) = 5.37, p = .02), chosen as more trustworthy by 58% (p = .04), preferred by 62% (p = .001), rated as based on more appropriate evidence (M = 0.83 against 0.79, F(1,182) = 13.68, p = .0003), and judged less mentally demanding (2.79 against 2.94, F(1,182) = 7.75, p = .0006), with 61% saying the deductive AI required more thinking (p = .005). Objective performance on the proxy task, however, was identical across explanation types (M = 0.64 both, F(1,182) = 0.0009, n.s.), including on error trials (0.40 against 0.41, n.s.).
On the actual decision task the subjective ordering reversed. Deductive explanations were now trusted more (M = 3.68 against 3.44, F(1,64) = 5.96, p = .01), with 65% naming them more trustworthy (p = .02), rated more helpful (3.92 against 3.65, F(1,64) = 3.66, p = .06), with 68% choosing them (p = .006), and preferred overall by 63% (p = .05). The same two explanation designs, evaluated by the same population on the same material, produced opposite preferences depending only on whether the task was to predict the AI or to make the decision.
Explanations did help in absolute terms. AI assistance of any kind raised accuracy from 0.46 to 0.72 (F(1,2446) = 118.07, p < .0001), and adding explanations raised it further from 0.68 to 0.74 (F(1,2014) = 5.10, p = .02). Explanations also raised trust (3.56 against 3.17, F(1,483) = 11.28, p = .0008), helpfulness (3.78 against 3.26, F(1,147) = 4.88, p = .03) and perceived understanding (3.84 against 3.67, F(1,2014) = 6.89, p = .009) relative to a bare recommendation.
The second dissociation is between preference and performance within the real task. Overall accuracy did not differ by explanation type (F(1,64) = 0.44, n.s.), but a strong interaction with AI correctness did (F(2,2013) = 15.03, p < .0001). When the AI was right the two were equivalent (inductive 0.78, deductive 0.81, n.s.). When the AI was wrong, participants seeing inductive explanations were substantially more accurate than those seeing deductive ones (0.63 against 0.48, F(1,64) = 7.02, p = .01). The explanation type participants trusted less, preferred less and found less helpful was the one that protected them from the AI's errors. Both experiments replicated.
The think-aloud study supplies the mechanism, and it is a difference in how the two explanations are used rather than in how much information they carry. Eight of eleven participants preferred inductive explanations, reading the four images as data backing the recommendation. More revealingly, with inductive explanations participants tended to form their own judgement first and then use the recommendation to confirm it, whereas with deductive explanations they evaluated the explanation itself. Because the erroneous examples were visible as images, an inductive explanation of a wrong recommendation exposed the error; a deductive list of misrecognised ingredients apparently did not trigger the same scrutiny.