← Back to the framework

You Complete Me: Human-AI Teams and Complementary Expertise

N = 518 Other

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Manipulated
LowHigh
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty Low (diagnostic)
AI Ecosystem Designable
AI performance 60%-70%
AI design
Protocol Human-first (update)
AI stance Directive
Interactivity Static
XAI Manipulated
Presence Manipulated
Type None · Narrative
Quality Correct
Number of AI advisors One
Research questions

Study 1: RQ1: How does the degree of complementary expertise affect team performance, reliance behavior, and trust in AI-assisted decision making?

Study 2: RQ2: How does embracing or distancing language in AI explanations affect trust, reliance behavior, and team performance in AI-assisted decision making across different levels of complementary expertise?

In the authors' words

"when the AI demonstrates its expertise (or lack thereof) in areas of the task that the participant has little to no expertise, the participant is able to discern the AI's (non-)expertise and adjust their reliance on its recommendations over time, instead of just blindly agreeing with the AI"

results, study 1

"Subjective trust may thus depend more strongly on complementary expertise than the type of explanation"

results, study 2 summary
Experimental design

Two between-subjects experiments on Amazon Mechanical Turk sharing one task, with a combined 518 valid participants.

Task. Object identification. Participants saw images of 2D shapes and assigned each to one of six categories. Three were Regular shapes every participant already knows: rectangle, triangle, circle, varying randomly in fill colour, side length and corner angle. Three were Fake shapes invented for the study and given nonsense names: Senectus, Pharetra, Ultrices, built from Bezier-curve sides with border and interior patterns of dots, dashes or both. Fake categories are distinguished by the number of sides and by whether the border and interior patterns match, while fill colour, side length and curve angle vary randomly as category-irrelevant noise.

The design's purpose is to make every participant a guaranteed total expert on half the stimuli and a guaranteed total non-expert on the other half, without any training, so that complementarity can be constructed rather than assumed.

AI partner. An assistant named ShapeBot was configured to have perfect expertise in exactly three of the six categories in every condition, so the amount of AI competence is constant and only its location moves. The four complementarity levels are: level 0, correct on all 3 Regular and none of the Fake (complete overlap with the human); level 1, 2 Regular and 1 Fake; level 2, 1 Regular and 2 Fake; level 3, 0 Regular and 3 Fake (perfect complementarity). Outside its assigned categories ShapeBot performed at chance, roughly one in three.

Protocol. Forty-two trials, seven per category, randomly ordered. Each trial ran: the participant makes an initial guess, then sees ShapeBot's recommendation, then makes a final decision, then receives feedback showing whether ShapeBot and the participant were correct and how many points were earned (one point for a correct first guess plus one for a correct final guess). Five training trials on six different shapes preceded the main task.

Study 1 (N = 160 of 178 recruited). Between-subjects with complementarity level as the only factor. Dependent variables: first guess performance, final decision performance, agreement frequency, switch-to-agree frequency, and subjective trust measured with the seven-item Muir and Moray Scale of Trust in Automated Systems covering competence, predictability, dependability, responsibility, reliability and faith.

Study 2 (N = 358 of 395 recruited). The same task with a written explanation added alongside the recommendation, in a factorial design crossing complementarity level (4) by point of view (2) by belief marker (2), plus a no-explanation control. Explanations followed a fixed template: [point of view] [belief marker] this is a [recommendation] because [point of view] [belief marker] it has [reasons]. Point of view was first person (I) or third person (ShapeBot); belief markers were embracing (knows, realizes) or distancing (thinks, believes). Critically, when ShapeBot recommended the wrong category the stated reason still accurately described the object shown, because the error was constructed as a classification error rather than a perception error.

Study 2 analysis split the data by whether ShapeBot was assigned to be correct on that shape category (Expert AI cases) or to perform at chance (Non-Expert AI cases), and restricted attention to Fake shapes, since study 1 showed participants ignored the AI on Regular shapes.

Full findings

Where the AI's competence sits relative to the human's determines team performance, even when the amount of AI competence is held constant. This is the paper's contribution and it is cleanly demonstrated: ShapeBot was perfectly expert on exactly three of six categories in every condition, so overall AI accuracy barely moved, yet final decision performance rose steeply and monotonically with complementarity (study 1: F(3,159) = 145.732, p < .001, partial eta squared = .737 across all shapes; F(3,159) = 150.537, p < .001, .743 on Fake shapes alone, with every pairwise comparison significant). Study 2 replicated it at F(3,357) = 306.537, p < .001, .722.

People worked this out for themselves, from behaviour alone, with no explanation in study 1. Agreement with the AI fell as complementarity rose when all shapes are pooled (F(3,159) = 32.900, p < .001, .388) but rose on the Fake shapes where help was actually needed (F(3,159) = 45.417, p < .001, .466), and switch-to-agree followed the same pattern (F(3,159) = 29.101, p < .001, .359 on Fake shapes). The learning is visible across trials: agreement started at the same level in all conditions and then diverged, climbing under perfect complementarity (beta = 0.117, p = .005) and falling under complete overlap (beta = -0.145, p < .001), interaction F(18,3359) = 2.958, p = .007. First guess performance was unaffected by condition (F(3,159) = 0.281, p = .839), confirming that the accuracy gain came from selective reliance rather than from participants learning the shapes.