Assertiveness-based Agent Communication for Personalized Medicine on Medical Imaging Diagnosis
N = 52 Medical imagingClinical assessment / diagnosisThis study in the framework
Held constant across all three trials. The same two models drove every condition: a DenseNet estimating lesion scores from 2D mammography and ultrasound images and a 3D ResNet for DCE-MRI volumes, each producing a separate per-modality prediction over three BIRADS classes rather than a fused multimodal one. Training used 289 of 338 acquired cases, roughly 2890 images, with BIRADS ground truth assigned by the head of radiology at one of the participating hospitals. The authors report model performance only as an improvement over their own previous system, an average decrease of about 26% in the false-positive rate and about 2% in the false-negative rate, and justify preferring false-positive and false-negative rates to accuracy because the dataset is imbalanced. Case-level information was displayed to clinicians in every condition: the agent reported the accuracy of the output, heatmap values and, in the abstract's terms, the sensitivity and specificity of the agent. Because the manipulation was purely one of tone, model behaviour is identical across conditions by construction.
RQ1. How does an assertiveness-based agent affect medical assessments? H1.1. Efficiency of clinicians in terms of time performance per each diagnosed patient will be higher with an assertiveness-based agent. H1.2. Classification accuracy of clinicians will not suffer with an assertiveness-based agent. H1.3. Through assertiveness-based communication, accuracy differences between novice and expert clinicians will depend on the tone of the personalized explanations.
RQ2. How is an assertiveness-based agent perceived by clinicians? H2.1. Clinicians will have a preference for an assertiveness-based agent. H2.2. Clinicians will consider an assertiveness-based agent more trustworthy. H2.3. Personalized highlights and explanations will not increase clinicians' workload nor decrease usability. H2.4. Novice and expert clinicians will perceive reliability and capability differently, depending on the levels of assertiveness.
"we show that personalizing assertiveness according to the professional experience of each clinician can reduce medical errors and increase satisfaction, bringing a novel perspective to the design of adaptive communication between intelligent agents and clinicians"
abstract
"the chance of a patient getting classified correctly by a novice was significantly higher (Accuracynovice = 91%) with the assertive agent (i.e., imposing AI recommendations) than with the non-assertive (i.e., more suggestive AI recommendations). On the contrary, the chance of correctly classifying the patient by an expert clinician was slightly higher (Accuracyexpert = 78%) with the non-assertive agent"
results, H1.3
"agents may need to be more assertive for novice clinicians, while suggestive tone may be more appropriate for expert clinicians"
results, H1.3
"clinicians took less 25% of the time to diagnose a patient with the assertiveness-based agent in comparison to the conventional agent"
discussion
Experimental design
A within-subjects, counterbalanced experiment with 52 clinicians, each completing three trials, preceded by 52 semi-structured interviews and user-centred design activities from which the research questions were derived. Participants were volunteers recruited from 11 Portuguese clinical institutions including public hospitals, cancer institutes and private clinics. Ethics approval was obtained from each institution.
Participants and expertise. Two professional experience groups were compared. Experts made up 55.77% of the sample: 34.62% seniors with more than ten years of practice and 21.15% middle clinicians with five to ten years. Novices made up 44.23%: 32.69% juniors with up to five years after the specialty exam and 11.54% interns who had not yet taken it.
Task. Breast cancer diagnosis from medical imaging. For each patient the clinician read roughly six views across three modalities (two CC and two MLO mammography views, one ultrasound, one DCE-MRI volume of 100 to 200 frames) and issued a BIRADS severity classification, the task ending when the clinician accepted or rejected the BIRADS proposed by the agent. Each clinician saw three patients per trial, drawn at random from a pool of 289 classified cases, selected to span three severity bands: BIRADS 1 (no findings), BIRADS 2-3 (benign or probably benign) and BIRADS 4-5 (suspicious or highly suspicious malignancy). Clinicians could request a visual explanation inside the image during the task.
Conditions. Three trials, presented counterbalanced, differing only in the tone and granularity of the same underlying model output. The conventional agent reported the suggested BIRADS, the model's accuracy and heatmap values as bare numeric estimates. The assertiveness-based agent added a bounding box on the image and a descriptive sentence of the clinical arguments, delivered in one of two registers: assertive, imposing the recommendation with wording such as "must", and non-assertive, suggesting it with wording such as "it looks like". Figure 1 of the paper contrasts the two. All clinicians experienced all three trials.
AI models. A DenseNet estimating lesion scores for 2D mammography and ultrasound and a 3D ResNet for MRI volumes, each producing a per-modality prediction over three BIRADS classes rather than a fused multimodal one. Both were trained on 289 of 338 acquired cases, roughly 2890 images, with ground truth assigned by the head of radiology at one institution.
Full findings
The same recommendation, said in a different register, produced different diagnoses: an imposing tone helped novices and a suggestive tone suited experts, and the assertiveness-based agent cut diagnosis time by a quarter without costing accuracy.
Efficiency (H1.1, supported). Clinicians diagnosed a patient in 124.02 seconds (SD 44.60) with the assertiveness-based agent against 166.12 seconds (SD 60.42) with the conventional one, F = 11.32, p = .005, r = .49, a large effect. The authors describe this as taking 25% less time.
Accuracy (H1.2, supported as a null). There was no significant difference in overall classification accuracy between agents, F = 1.85, p = .37. The hypothesis was framed as a non-inferiority claim, so the null is the intended result: richer, more assertive communication did not degrade accuracy.
Tone by expertise (H1.3, supported, and the paper's substantive finding). The association between level of assertiveness and professional experience was significant, chi-square = 3.84, p = .001. Novices classified patients correctly more often with the assertive agent (accuracy 91%), while experts did slightly better with the non-assertive one (accuracy 78%). In odds terms, novices were 17.4% more likely to be correct with the assertive agent, experts 4.4% more likely with the non-assertive agent.
The decision table shows where the gain came from. For novices, overall correct decisions rose from 69.70% with the conventional agent to 81.59% with the assertive agent and 75.63% with the non-assertive one; overall mistakes fell from 30.30% to 18.41%. Almost all of that movement is in wrong rejects, which fell from 26.6% to 15.36%, meaning novices stopped overturning correct AI recommendations. Wrong accepts barely moved, 3.7% to 3.05%, so the assertive tone did not buy the accuracy gain with over-reliance. Experts changed little on any measure: 63.77% conventional, 65.76% assertive, 66.41% non-assertive, with wrong accepts drifting up slightly under the assertive tone (4.1% to 4.75%).
Preference (H2.1, supported). Preference for the assertiveness-based agent was significant across the four experience bands, F = 8.35, p = .001, r = .41. Of the 52 participants, 66% preferred the assertiveness-based agent, 10% expressed no preference.
Trust (H2.2, partly supported). Overall trust did not differ significantly, F = 19.47, p = .06, and neither did perceived understanding, p = .14. The two components that did move were competence, p = .04, and thoughtfulness, p = .001, both favouring the assertiveness-based agent.
Workload and usability (H2.3, supported as a null). No significant difference in NASA-TLX workload, p = .38, or SUS usability, p = .38. The added explanatory content did not cost mental effort.
Perceived reliability and capability (H2.4, supported). Both differed significantly by assertiveness level, reliability F = 31.36, p = .0001, capability F = 18.17, p = .0003.
The mechanism here is communicative rather than informational. Across all three trials the underlying model, its predictions and the clinical facts conveyed were the same; what changed was the register in which the recommendation was asserted, and the optimal register depended on the recipient's experience.