EN /PT

AI-Assisted Decision-Making Framework

A framework for AI-assisted decision-making, built from an evidence base of scientific studies. Every AI-assisted decision has three components: the Human and the Task, which characterize each decision setting, and the AI Assistant ecosystem, the only component that can be leveraged to improve the collaboration. It's public, it's a working draft, and it changes as the evidence base grows.

Version V1.0.0 · Updated 2026-08-20 · CC-BY ·

AI-assisted decision-making

What it is

AI-assisted decision-making describes any process where a person makes the final call but an AI system shapes it along the way, sometimes with an explicit recommendation, sometimes with something softer: a risk score, a highlighted region of an image, a short explanation. A radiologist reads a mammogram while a model flags suspicious tissue. A credit analyst sees a predicted probability that an applicant will default. A judge is shown a recidivism risk score.

The appeal is that the two should cover each other's weaknesses: people are inconsistent and carry biases, models are consistent but brittle outside the cases they were trained on. The evidence for that promise is thinner than it sounds. Human-AI teams usually beat people working alone, but they often fall short of the AI working alone.

The reason is reliance. Over-reliance means accepting a recommendation that is wrong; under-reliance means rejecting one that is right. Either is enough to undo the benefit, and neither is fixed simply by showing the person more information about the model.

This has made AI-assisted decision-making an intensively studied problem, and a large body of experiments now tests interventions (explanation designs, confidence displays, protocols that withhold the recommendation until the person has committed to their own judgement) to work out which techniques and designs genuinely improve the collaboration, curb over- and under-reliance, and raise the quality of the decision.

Why it matters

Since 2 August 2026, this is no longer only a research question. Article 14 of the EU AI Act obliges providers of high-risk AI systems (a category that covers medical devices, creditworthiness assessment, recruitment, education and law enforcement) to design them so that they can be "effectively overseen by natural persons" while in use. The Act is unusually specific about what that oversight has to withstand. Among the capacities the system must enable is:

“to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions to be taken by natural persons.”

Regulation (EU) 2024/1689, Article 14(4)(b)

The regulation names the precise failure this literature has spent a decade measuring. It names the other side too: Article 14(4)(d) requires that the overseer be able to "disregard, override or reverse" the output, the capacity whose overuse is under-reliance.

What the law does not say is how. "Remain aware" is a goal, not a mechanism, and the obvious routes to awareness turn out not to work reliably. Explanations frequently fail to reduce over-reliance and sometimes increase it. Interventions that do reduce it can cost accuracy, speed, or the user's willingness to adopt the system at all. A legal duty to counter automation bias is only actionable if we know which design choices actually move it, for whom, and on which tasks, and that is an empirical question about human behaviour, not a question about the model.

The framework

This framework is an attempt to put that evidence in order. It collects experimental studies of AI-assisted decision making, codes each one the same way, and asks what is actually known about how people behave when they decide alongside a machine, for researchers who need to see where the evidence is thin, for practitioners designing these systems, and for anyone who wants the picture without reading a hundred papers.

It divides the setting into three components. The human is who decides: their expertise, whether they work alone or in a group. The task is what is being decided: how hard it is, what it costs to be wrong, whether there is time to think, whether the answer can be verified at all. Both are largely inherited (a hospital cannot staff radiology with non-radiologists). The AI ecosystem, the model's accuracy, how its output is presented, whether it explains itself, is the part that is designed, and where interventions live. Within each sit the factors that shift decision performance and reliance.

It does not stop at organising the literature. Each factor carries mechanisms: short, directional claims about what causes what, linked to the studies that support them and the studies that contradict them. Mechanisms are ordered by how much support they have, from the most-replicated finding down to a single study. The aim is a resource that answers where the evidence permits an answer, and says so plainly where it does not.

Human Inherited
Task Inherited
AI Ecosystem Designable

Every framework is a simplification, and this one is no exception. The factors shown here are deliberately coarse; individual studies contain distinctions, conditions and caveats that no diagram can hold. Where a study is richer than its position in the framework suggests, that detail sits on the study's own page.

Where the evidence stands

22 mechanisms, ordered by how many studies support each, from the mechanism with the highest number of supporting studies down to the one resting on the fewest, currently 2 studies.

The effect of explanations on decision quality is conditional, not a main effect, and is often null or negative

11 supporting ·
Decision accuracyTrustOver-reliance
Details

Asked unconditionally, do explanations help, this literature answers no. Asked conditionally, specifying model performance, task difficulty, risk and explanation form, it answers yes under stated conditions and no otherwise.

  • On the negative side, Alufaisan and colleagues ran the three-arm design with a no-AI control, an AI-only arm and an AI-plus-explanation arm, and found the AI itself carried the benefit while explanations added nothing.
  • De Brito Duarte and colleagues found the effect of explanation on trust jointly conditional on system performance and risk level, and only feature importance moved anything, with counterfactuals indistinguishable from no explanation for lay users.
  • On the positive side, Vasconcelos and colleagues obtain clear benefits in hard tasks and with salient explanations, and Leichtmann and colleagues find explanations improving accuracy on precisely the items where the classifier erred, while lowering trust in a way they read as better calibration.

Calibrated reliance requires knowing when the AI fails, which case-level confidence supports and explanations do not

9 supporting ·
Over-relianceUnder-relianceDecision accuracy
Details

Knowing that a model is 80% accurate tells the decision-maker nothing about the case in front of them. What supports selective reliance is a usable model of the error boundary, meaning where in the input space the system fails, and the properties of that boundary matter as much as the accuracy figure attached to it.

The two Bansal 2019 papers establish both halves. Mental model quality mediates the relationship between AI accuracy and team performance, and boundaries that are less parsimonious, higher-dimensional or stochastic are learned poorly or not at all: a two-sided stochastic boundary was never learned, with the proportion of optimal choices flat near 50% across rounds. The corollary is that an update improving accuracy can reduce team performance by invalidating the boundary the human has already learned.

Cutting over-reliance usually raises under-reliance, so accuracy does not move

8 supporting ·
Over-relianceUnder-relianceDecision accuracy
Details

Most interventions that make people less willing to accept the AI make them less willing to accept it when it is right as well as when it is wrong. The two movements cancel, which is why so many designs succeed on one reliance measure and vanish on decision quality. A reported reduction in over-reliance is therefore uninformative unless the corresponding under-reliance figure is given alongside it.

Explanations help only when the user can check them against the case

8 supporting ·
Over-relianceDecision accuracyUnderstanding of AI
Details

What determines whether an explanation improves a decision is not how much it explains but whether the user can hold it against the evidence and see a mismatch. Explanations that are inspectable at a glance, such as similar cases shown as images or a highlight drawn on the artefact itself, let people form their own reading first and make the model's errors visible. Explanations that must be reasoned about in the abstract, such as feature weights, counterfactuals or fluent prose, are rhetorically dominant, suppress the user's own reading of the case, and are accepted without being checked.

Fluency makes this worse: the most conversational interface in the base produced the worst self-reliance. The mechanism is distinct from explanation quality, since a faithful explanation in an uncheckable form still displaces judgement.

About the evidence base

I didn't start with a framework. I started with papers: experimental studies of people making decisions with AI. What became obvious quickly was that the same handful of variables kept reappearing: who the decision-maker was, what kind of task they faced, how the AI presented itself and how good it was. Different labs, different domains, different vocabulary, but a recurring set of things being varied and measured. Those recurrences became the framework's nine factors, grouped into the three components that seemed to organise them, the human, the task, and the AI ecosystem. The factors came out of the papers rather than out of theory, which is a strength and a limitation at once: they describe what this literature has chosen to study, not necessarily what matters most about human-AI decision-making.

Once the factors existed I went back and coded every study against the same fixed schema, recording what each one manipulated, what it held constant, and what it simply didn't report. A shared schema makes studies comparable that were never designed to be compared, and it makes disagreement visible. From there I could read off claims (e.g. this factor moves that outcome, under these conditions) that more than one study speaks to. Those are the mechanisms, and each one carries both the studies that support it and the studies that contradict it.

The coding is open. Every study sits in the Framework Evidence Base with its full codification and the prose behind each judgement, and every code is defined in the Codebook. Any claim on this site can be traced back to a paper.

This framework covers the studies I know about. It is not a systematic review, and it is certainly incomplete. If you know of experimental work that belongs here, especially work that contradicts something on this page, please send it to me and I will code it in.

This framework is my own work. I built it, and I am responsible for its content and for any errors in it.

I used AI assistance in the process: to help edit text, to extract study details into the common coding scheme, and to propose links between studies and mechanisms. Every paper in the evidence base was read by me, every coded field was checked by me, and every mechanism and link was my decision.