cs.AIOct 7, 2026

Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study

Authors: Lin Wu, Zhe Xu, Hongyi Wang, Feifei Zhou, Wei Deng, Chunlong Zhang, Yuting Zhu, Kaixiao Chen, +5 more

Organizations: Department of Radiology, The First Affiliated Hospital, Jiangxi Medical College, Nanchang University · Jiangxi Province Medical Imaging Research Institute · Department of Computer Science and Engineering, Hong Kong University of Science and Technology, Hong Kong, China · Department of Radiology, Shangrao City People’s Hospital, Shangrao, China · Department of Radiology, Nanchang People’s Hospital, Nanchang, China · Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology, Hong Kong SAR, China · Division of Life Science, Hong Kong University of Science and Technology, Hong Kong SAR, China · State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology, Hong Kong SAR, China · HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology, Futian, Shenzhen, China

Abstract

Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.

Figures & tables

Explore similar work

CardsList
  1. Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

    Sep 23, 2026Jiaju Huang, Hao Yang, Xinyu Ma +6Radiology Report GenerationClinician Trust

  2. Vision-language models for chest radiography do not always need the image

    Jun 16, 2026Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams +4Chest X-RayMedical Vision-Language Models

  3. RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

    Apr 16, 2026Mélanie Roschewitz, Kenneth Styppa, Yitian Tao +10Radiology Report GenerationMedical Vision-Language Models