← Back to the framework

AI, Help Me Think—but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support

N = 21 ETF Investement

This study in the framework

Human Inherited
Expertise level Lay users
Team composition Individual
Task Inherited
Difficulty Not reported
Stakes Not reported
Stress level / time constraint Not reported
Task uncertainty High (prognostic)
AI Ecosystem Designable
AI performance Not applicable
AI design Manipulated
Protocol Manipulated (AI-first, Human-first (update))
AI stance Manipulated (Directive, Reflective)
Interactivity Static
XAI Not reported
Number of AI advisors One
Research questions

RQ: How do different paradigms for AI decision support, one that extends users' reasoning and one that provides direct recommendations, affect users' decision-making processes, perceptions of the AI, and decision outcomes in complex decisions?

In the authors' words

"we intentionally chose an open-ended task without any objectively correct or false decisions. While this reflects many real-world tasks... it also meant that it was difficult to establish clear performance metrics."

limitations

"I think it gives me a bit more control and agency... It didn't just tell me, 'This is what you should do.'"

findings, 6.2.1, participant RE-12

"It allows me to delve deeper into the problem and understand the recommendations more than blindly trusting."

findings, 6.2.2, participant ER-9

"I hadn't taken any decisions. Then I was trying to make my mind based from what the assistant tells me."

findings, 6.2.3, participant ER-4
Experimental design

Within-subjects study with two AI conditions in randomised order, preceded by an unaided familiarisation step.

Task. Building and revising an ETF portfolio across three simulated time steps, August 2024, 2026 and 2028, starting with USD 10,000 and receiving a further USD 10,000 at each step. The first step was familiarisation with no AI. The second and third used the two AI systems in randomised order. The instrument set was 31 equity ETFs.

Conditions. RecommendAI gave direct ETF recommendations without requiring anything from the user first. ExtendAI required the user to write a rationale for their intended changes, then returned feedback embedded in that writing, and was instructed never to name specific ETFs. Both were GPT-4 via the API, given identical ETF data, identical investor-profile information and otherwise identical instructions, so the contrast is the paradigm rather than the underlying model.

Measures. Five-point questionnaires on perceived informatedness, confidence and satisfaction, administered before and after portfolio outcomes were revealed; a shortened NASA-TLX for cognitive load; an agency slider running from −50, the AI decided, to +50, I decided; portfolio diversification metrics covering the number of countries and the balance across countries and sectors; two proxies for AI influence, the percentage of trades deviating from the rationale the user had written under ExtendAI and the percentage of trades matching the recommendations under RecommendAI; time spent and the number of ETFs examined; and exit interviews averaging 23 minutes (SD = 6.8).

Full findings

Making the AI respond to the user's own reasoning rather than hand them an answer roughly halved its behavioural pull and doubled the time and breadth of the user's own search, at the cost of actionability and novelty.

The influence proxies are the clearest quantitative result. Under RecommendAI 45.00% of trades matched the AI's recommendations; under ExtendAI 23.08% of trades departed from the rationale the participant had written themselves. Engagement moved with it: participants spent 17.46 minutes (SD = 6.03) in the ExtendAI condition against 8.56 (SD = 3.37) with RecommendAI, of which 8.38 minutes (SD = 4.69) went on writing the rationale, and they examined 19.67 of the 31 available ETFs (SD = 9.23) against 12.86 (SD = 8.79). Extending reasoning is slower and broader; recommending is faster and narrower.

Perceptions split along the same seam and, importantly, not in the same direction as behaviour. RecommendAI produced higher confidence (86% against 67%) and more reported new insights (67% against 52%), while ExtendAI produced higher satisfaction (67% against 43%), higher trust (76% against 71%), more consideration of the AI's input (90% against 81%) and higher rated helpfulness (81% against 71%), at slightly greater cognitive load (NASA-TLX M = 57.00, SD = 14.02 against M = 52.52, SD = 15.55). Both systems raised felt informedness equally, from 57% unaided to 71%. Preference was almost exactly split, 11 for RecommendAI and 10 for ExtendAI, and 86% said they would use RecommendAI in the real world against 76% for ExtendAI. Diversification improved under both, marginally more under ExtendAI on country count and country balance.

The authors organise the discussion around three tensions rather than a winner. Actionability against cognitive engagement: a named ETF is easy to act on and cheap to think about, while general feedback preserves agency but leaves the user to work out what to do. Novelty against alignment: RecommendAI produced out-of-the-box suggestions that were hard to verify, while ExtendAI's feedback sat inside the user's own reasoning and so was easy to verify but rarely surprising. Timing: RecommendAI arrives before deliberation and risks anchoring, ExtendAI arrives after it, which helps someone still forming a view but does little for someone whose rationale is already settled, where the authors observed confirmation bias instead.

The claim the authors make for ExtendAI is about trust calibration rather than accuracy: because the feedback is embedded at a fine grain in the user's own writing, it is easier to check, and confidence tracked satisfaction rather than running ahead of it, which they read as warranted trust rather than over-reliance.