Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making
N = 600 Biography classificationThis study in the framework
Held constant by construction, and identical across all three conditions. Two logistic regression models were trained on mutually disjoint vocabularies, one task-relevant and one gendered, and the 14 bios shown to participants were selected so that both models produce the same prediction with similar probabilities. Only the explanation differs between conditions; the recommendation never does. What participants experienced. Of the 14 bios, 8 carry a correct AI recommendation (three WTT, three MPP, one WPP, one MTT) and 6 carry an incorrect one (three WPT, three MTP), giving an experienced accuracy of 57.1%. This is the figure the row is binned on. No hold-out accuracy for either classifier is reported.
RQ: "we study the effects of feature-based explanations on distributive fairness of AI-assisted decisions, specifically focusing on the task of predicting occupations from short textual bios. We also investigate how any effects are mediated by humans' fairness perceptions and their reliance on AI recommendations."
"we see that such explanations do not enable humans to discern correct and incorrect AI recommendations. Instead, we show that they may affect reliance irrespective of the correctness of AI recommendations."
abstract
"while disparities in error types decreased in the gendered condition compared to the baseline, this was due to a shift in the types of errors, as opposed to an increased ability to override mistaken AI recommendations."
results, distributive fairness
"the gendered condition fostered more overrides of AI recommendations when a woman was predicted to be a teacher, irrespective of whether this prediction was correct; meanwhile, in the task-relevant condition participants were less likely to override recommendations where a man was predicted to be a professor, irrespective of his true occupation."
discussion, summary of findings
"this means that explanations may not only be exploited to mislead people's fairness perceptions but also their reliance behavior."
results, role of fairness perceptions
Experimental design
A randomised between-subjects online experiment with three conditions, cleared by an institutional ethics committee.
Task. Predicting from a short textual bio whether the person is a professor or a teacher, with an AI recommendation shown. The domain is framed as AI-assisted hiring, where occupation must often be inferred from unstructured online text, and where misclassification systematically excludes people from exposure to opportunities. Professor is historically a men-dominated occupation and teacher is associated with women, so the model's errors fall along stereotype lines.
Dataset. The public BIOS dataset of roughly 400,000 online bios across 28 occupations, restricted to professors and teachers, leaving 134,436 bios of which 118,215 are professors and 16,221 teachers. Gender is inferred from pronouns, which excludes non-binary people, a limitation the authors state. Within the restricted set, 55% of professor bios are men and 60% of teacher bios are women.
Conditions. Baseline shows the bio and the AI recommendation with no explanation and no colour coding. Task-relevant shows the same recommendation with LIME word highlighting from a model trained on a task-relevant vocabulary. Gendered shows the same recommendation with LIME highlighting from a model trained on a gendered vocabulary. Colour indicates which occupation a word supports and colour intensity indicates its weight.
The two classifiers. The design's central device is two logistic regressions trained on mutually disjoint vocabularies. The task-relevant vocabulary contains words appearing more often, for both men and women, in professor or teacher bios than in any of the other 26 occupations, such as faculty, kindergarten and phd. The gendered vocabulary contains the words most predictive of gender, including pronouns and husband and wife but also dance, art and engineering, which are not evidently gendered yet highly correlated with the sensitive attribute. Logistic regression was chosen specifically so that the LIME explanations would be faithful to the underlying model. Crucially, the two models agree on all 14 bios shown, so the recommendation is identical across all three conditions and only the explanation differs.
Stimulus construction. 14 bios, structured by gender of the person, true occupation, and AI recommendation. Three each of correctly recommended women teachers (WTT), incorrectly recommended women professors (WPT), incorrectly recommended men teachers (MTP) and correctly recommended men professors (MPP), plus one correctly recommended woman professor (WPP) and one correctly recommended man teacher (MTT). The focus is deliberately on errors that align with gender stereotypes, since those are the errors an occupation model with gender bias actually makes. The WPP and MTT cases exist to pre-empt the inference that the AI always predicts teacher for women, and order was randomised subject to those two appearing among the first five, so that participants would not meet a run of stereotype-aligned errors early and have their reliance depressed. Bios were drawn from a random holdout set, matched on length, required to receive the same prediction and similar probabilities from both classifiers, and required not to have probabilities that were too high, so as to exclude cases that were too easy; the authors then screened the remainder manually.
Procedure. Participants were first asked what they take the difference between professor and teacher to be. They then completed the 14 bios one at a time. Those in the two explanation conditions answered three fairness-perception items afterwards, deliberately placed after the task so that eliciting perceptions could not moderate reliance. The baseline received no perception items, since it offers no cue about the AI's procedure. An open-ended question then asked what information they had relied on, which the authors used to confirm that the professor-teacher distinction was construed consistently across conditions. A demographic survey closed the study.
Measures. Reliance is decomposed into four cells: correct adherence, detrimental adherence, corrective overriding and detrimental overriding, with the first and third summing to final decision accuracy. Distributive fairness is measured as absolute gender disparity in error rates, separately for promoting errors (teacher predicted professor) and demoting errors (professor predicted teacher). Fairness perceptions are the average of three 5-point items, two adapted from Colquitt and Rodell's procedural justice construct and one written for this setup about whether it is fair for the AI to consider the highlighted words, with the third asked last and non-revisable to avoid priming; scale reliability was Cronbach's alpha 0.77.
Analysis. Nonparametric throughout, because normality and equal variance could not be confirmed: Kruskal-Wallis omnibus tests and two-tailed Mann-Whitney U tests for pairwise comparisons, with 0.05 < p < 0.1 reported as marginal. OLS regressions relate overriding to fairness perceptions.
Full findings
Explanations changed which errors people made without changing how many. Distributive fairness moved in both directions depending on what the explanation highlighted, but never because participants got better at catching the AI's mistakes.
Accuracy did not move. Mean accuracy was 59.49% (SD 13.11) in the baseline, 56.94% (SD 13.86) with task-relevant explanations and 57.96% (SD 14.30) with gendered explanations, with no significant difference (Kruskal-Wallis, p = 0.260), and no difference for men bios (p = 0.199) or women bios (p = 0.151) separately. Participants were paid per correct answer, so this is not an attention artefact.
Overriding did move. Participants in the gendered condition overrode more than in the task-relevant condition (p = 0.005), with detrimental overrides significantly above baseline (p = 0.012). The task-relevant condition produced the fewest overrides overall, with corrective overrides marginally below baseline (p = 0.097). The decisive negative result is that the ratio of corrective to detrimental overrides was unchanged by either explanation, at the aggregate level and separately for men and women bios. Explanations shifted the volume of overriding, not its aim.
What the overrides were aimed at instead was the stereotype. In the gendered condition participants performed marginally more corrective overrides for women (p = 0.083) and unchanged numbers for men (p = 0.588); in the task-relevant condition they performed fewer corrective overrides for men (p = 0.011) and unchanged numbers for women (p = 0.834). Because stereotype-aligned detrimental overrides did not differ across conditions, the authors infer that the gendered condition increased stereotype-countering detrimental overrides too. Their summary is precise: in the gendered condition participants became more likely to override a recommendation that a woman is a teacher "irrespective of her true occupation", and in the task-relevant condition less likely to override a recommendation that a man is a professor, again irrespective of the truth.
The fairness consequence follows arithmetically. In the baseline, participants already promoted men more than women (58.9% against 39.9%) and demoted women more than men (41.3% against 21.9%), giving absolute disparities of 19.0% for promotions and 19.3% for demotions. Task-relevant explanations pushed both further apart, hindering distributive fairness. Gendered explanations pulled both together, most sharply for demotions, which fell from 19.3% to 9.7%. The authors are careful about why: "while disparities in error types decreased in the gendered condition compared to the baseline, this was due to a shift in the types of errors, as opposed to an increased ability to override mistaken AI recommendations".
Perceptions carry the mechanism. Fairness perceptions were 3.53 (SD 0.85) in the task-relevant condition against 2.54 (SD 0.98) in the gendered condition, and perceptions related strongly and negatively to overriding (p = 1.10 x 10-11): participants overrode 52% of recommendations at the lowest perception level and 31% at the highest. But corrective and detrimental overrides rose at approximately equal rates as perceptions fell (corrective p = 9.18 x 10-5, detrimental p = 1.53 x 10-4), so the corrective-to-detrimental ratio is flat across perceptions. Perceived unfairness buys quantity of scepticism, not quality.
Two implications the authors draw are worth carrying into the framework. First, this extends algorithm aversion theory: perceived unfairness joins performance and agency as a driver of aversion. Second, and more troubling, because explanations can be engineered to hide the use of sensitive features, they can be used to manipulate not only perceptions, in the sense of fairwashing, but reliance behaviour itself. A system that conceals its use of proxies would produce higher fairness perceptions, fewer overrides, and worse distributive fairness, all at once.