On the Effect of Information Asymmetry in Human-AI Teams
N = 101 House price estimationThis study in the framework
What matters here is not the model's error but what was withheld from it. The property photograph was deliberately excluded from training, which is the entire basis of the study: the model's information set is a strict subset of the human's in one arm and identical to it in the other. The AI is therefore not weak in any absolute sense, it is blind to one channel.
H1: We hypothesize that a source of CP emerges from unique human contextual information (UHCI)", where CP is complementarity potential, the prerequisite the authors argue must exist before complementary team performance can be reached at all.
"CTP is rarely demonstrated in previous work as often the focus is on the design of explainability, while a fundamental prerequisite—the presence of complementarity potential between humans and AI—is often neglected."
abstract
"in order for human-AI decision-making to result in CTP, a more fundamental prerequisite is the presence of sufficient complementarity potential (CP) between humans and AI. In this context, we hypothesize that a source of CP emerges from unique human contextual information (UHCI)."
introduction
"In practice, domain experts often have access to further information not available to the AI during training as not all data might be digitally available due to technical or economic reasons."
introduction
"we find that in the presence of UHCI, humans become capable of positively adjusting the AI predictions resulting in CTP."
results and discussion
Experimental design
A single between-subjects online experiment with two treatments, run on Prolific.
Task. Real estate appraisal. Participants estimated the sale price of residential properties drawn from a public Kaggle dataset of Southern California house prices with accompanying photographs, 15,474 instances split 80/20 into training and test, with a hold-out set of 15 properties used as the experimental stimuli. Every participant saw all 15.
AI system. A random forest regression trained only on tabular features: street, city, number of bedrooms, number of bathrooms and square footage. The property photograph was deliberately withheld from the model during training. Uncertainty was derived from the individual trees of the forest as a predictive distribution, and the 5% and 95% quantiles were displayed to participants alongside the point prediction.
Manipulation. Whether the participant could see the property photograph. In the no UHCI treatment participants saw the same five tabular fields the model was trained on, so their information set matched the model's exactly. In the UHCI treatment they additionally saw the photograph, giving them unique human contextual information the model could not have used. This is the whole of the manipulation: the model, its prediction, its uncertainty display and the 15 properties were identical across arms.
Procedure. Both arms received an in-depth introduction to the dataset and the task, including summary statistics on property prices, followed by a comprehension check. Participants were explicitly told that the AI did not have access to the image during training, so the asymmetry was made salient rather than left to be discovered. For each of the 15 properties the sequence was fixed: make an independent price estimate first, which the authors state was done specifically to prevent participants entering a state of low cognitive activation; then see the AI's prediction with its confidence interval; then adjust the AI's prediction as well as possible. A demographics questionnaire followed all 15 instances.
Incentives. A base payment of 5 pounds, with the best-performing 10% of participants receiving an additional pound. The task took approximately 30 minutes.
Measures and analysis. Mean absolute error of the price estimate, computed separately for the participant's independent estimate and for the post-adjustment human-AI team output, and compared against the AI's own MAE on the same hold-out set. Student's t-tests with Bonferroni correction, with prerequisites verified in advance.
Full findings
Complementary team performance was achieved, and only in the arm where the human held information the model did not. This is the paper's single result, and it is offered as an existence proof rather than as a general effect.
Unassisted, the photograph was worth a great deal. Participants working alone reached a mean absolute error of $251,282 without the image and $200,510 with it, a difference of $50,772 (t = 4.6118, p < 0.001). The image is genuine signal, not decoration.
After adjusting the AI's prediction, the no UHCI team reached an MAE of $160,095 and the UHCI team $148,009, an improvement of $12,086 that is significant at the 0.05 level (t = 2.9571, p = 0.0155). The comparison that matters is against the model alone, which scored $163,080. The UHCI team beat the AI significantly (t = -4.6798, p < 0.001); the no UHCI team did not (t = -1.1596, p = 0.99). Where the human and the model saw exactly the same inputs, the team merely matched the model. Where the human saw one thing the model could not, the team exceeded it.
The effect is also much smaller than the unaided gap. The photograph was worth roughly $51,000 to a person deciding alone but only about $12,000 to the team, which suggests participants transferred only part of their private information into the adjustment. The authors do not analyse this shortfall, and they report no measure of how far participants moved from the AI's anchor, so the mechanism connecting private information to adjustment behaviour is not observed here.