You Complete Me: Human-AI Teams and Complementary Expertise
N = 518 OtherThis study in the framework
Manipulated, but through the location of competence rather than its amount. In every condition ShapeBot was perfectly accurate on exactly three of the six shape categories and performed at chance, roughly one in three, on the other three. Overall accuracy is therefore close to constant across conditions; what the four complementarity levels move is which categories the competence covers. Level 0 gives ShapeBot the three Regular shapes and none of the Fake ones, so its expertise completely overlaps the participant's and it is useless. Level 1 gives it two Regular and one Fake, level 2 one Regular and two Fake, and level 3 none of the Regular and all three Fake, so that its competence perfectly fills the participant's gap. Human expertise is fixed by construction at ceiling on Regular shapes and chance on Fake shapes, so complementarity is defined relative to a known human error boundary rather than estimated.
Study 1: RQ1: How does the degree of complementary expertise affect team performance, reliance behavior, and trust in AI-assisted decision making?
Study 2: RQ2: How does embracing or distancing language in AI explanations affect trust, reliance behavior, and team performance in AI-assisted decision making across different levels of complementary expertise?
"when the AI demonstrates its expertise (or lack thereof) in areas of the task that the participant has little to no expertise, the participant is able to discern the AI's (non-)expertise and adjust their reliance on its recommendations over time, instead of just blindly agreeing with the AI"
results, study 1
"Subjective trust may thus depend more strongly on complementary expertise than the type of explanation"
results, study 2 summary
Experimental design
Two between-subjects experiments on Amazon Mechanical Turk sharing one task, with a combined 518 valid participants.
Task. Object identification. Participants saw images of 2D shapes and assigned each to one of six categories. Three were Regular shapes every participant already knows: rectangle, triangle, circle, varying randomly in fill colour, side length and corner angle. Three were Fake shapes invented for the study and given nonsense names: Senectus, Pharetra, Ultrices, built from Bezier-curve sides with border and interior patterns of dots, dashes or both. Fake categories are distinguished by the number of sides and by whether the border and interior patterns match, while fill colour, side length and curve angle vary randomly as category-irrelevant noise.
The design's purpose is to make every participant a guaranteed total expert on half the stimuli and a guaranteed total non-expert on the other half, without any training, so that complementarity can be constructed rather than assumed.
AI partner. An assistant named ShapeBot was configured to have perfect expertise in exactly three of the six categories in every condition, so the amount of AI competence is constant and only its location moves. The four complementarity levels are: level 0, correct on all 3 Regular and none of the Fake (complete overlap with the human); level 1, 2 Regular and 1 Fake; level 2, 1 Regular and 2 Fake; level 3, 0 Regular and 3 Fake (perfect complementarity). Outside its assigned categories ShapeBot performed at chance, roughly one in three.
Protocol. Forty-two trials, seven per category, randomly ordered. Each trial ran: the participant makes an initial guess, then sees ShapeBot's recommendation, then makes a final decision, then receives feedback showing whether ShapeBot and the participant were correct and how many points were earned (one point for a correct first guess plus one for a correct final guess). Five training trials on six different shapes preceded the main task.
Study 1 (N = 160 of 178 recruited). Between-subjects with complementarity level as the only factor. Dependent variables: first guess performance, final decision performance, agreement frequency, switch-to-agree frequency, and subjective trust measured with the seven-item Muir and Moray Scale of Trust in Automated Systems covering competence, predictability, dependability, responsibility, reliability and faith.
Study 2 (N = 358 of 395 recruited). The same task with a written explanation added alongside the recommendation, in a factorial design crossing complementarity level (4) by point of view (2) by belief marker (2), plus a no-explanation control. Explanations followed a fixed template: [point of view] [belief marker] this is a [recommendation] because [point of view] [belief marker] it has [reasons]. Point of view was first person (I) or third person (ShapeBot); belief markers were embracing (knows, realizes) or distancing (thinks, believes). Critically, when ShapeBot recommended the wrong category the stated reason still accurately described the object shown, because the error was constructed as a classification error rather than a perception error.
Study 2 analysis split the data by whether ShapeBot was assigned to be correct on that shape category (Expert AI cases) or to perform at chance (Non-Expert AI cases), and restricted attention to Fake shapes, since study 1 showed participants ignored the AI on Regular shapes.
Full findings
Where the AI's competence sits relative to the human's determines team performance, even when the amount of AI competence is held constant. This is the paper's contribution and it is cleanly demonstrated: ShapeBot was perfectly expert on exactly three of six categories in every condition, so overall AI accuracy barely moved, yet final decision performance rose steeply and monotonically with complementarity (study 1: F(3,159) = 145.732, p < .001, partial eta squared = .737 across all shapes; F(3,159) = 150.537, p < .001, .743 on Fake shapes alone, with every pairwise comparison significant). Study 2 replicated it at F(3,357) = 306.537, p < .001, .722.
People worked this out for themselves, from behaviour alone, with no explanation in study 1. Agreement with the AI fell as complementarity rose when all shapes are pooled (F(3,159) = 32.900, p < .001, .388) but rose on the Fake shapes where help was actually needed (F(3,159) = 45.417, p < .001, .466), and switch-to-agree followed the same pattern (F(3,159) = 29.101, p < .001, .359 on Fake shapes). The learning is visible across trials: agreement started at the same level in all conditions and then diverged, climbing under perfect complementarity (beta = 0.117, p = .005) and falling under complete overlap (beta = -0.145, p < .001), interaction F(18,3359) = 2.958, p = .007. First guess performance was unaffected by condition (F(3,159) = 0.281, p = .839), confirming that the accuracy gain came from selective reliance rather than from participants learning the shapes.