cs.CVFeb 1, 2026

MedAD-R1: Consistency-Reinforced Policy Optimization for Interpretable Medical Anomaly Detection

Authors: Haitao Zhang, Yingying Wang, Jiaxiang Wang, Haote Xu, Hongyang Zhang, Yirong Chen, Yue Huang, Xinghao Ding

Organizations: School of Informatics, Xiamen University · Institute of Artificial Intelligence, Xiamen University · Zhejiang Expressway Co., Ltd. · School of Transportation Sclence and Engineering, Beihang University · School of Science and Engineering, Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory

Abstract

Medical Anomaly Detection (MedAD) offers a promising direction for medical image analysis with Large Multimodal Models (LMMs). However, progress is limited by fragmented datasets and the tendency of Supervised Fine-Tuning (SFT) to learn superficial image-text correlations rather than verifiable diagnostic reasoning. Consequently, current models often generate fluent explanations that are either insufficiently grounded in the image or inconsistent with their final answers, limiting their reliability in high-stakes medical applications. To address these issues, we introduce MedAD-38K, a large-scale, multimodal, and multicenter benchmark containing structured Visual Question Answering pairs and quality-controlled diagnostic Chain-of-Thought annotations across five core MedAD tasks. Based on this benchmark, we propose a two-stage framework. Cognitive Injection first uses SFT to inject domain-specific medical knowledge and establish a structured think-then-answer format. Consistency Group Relative Policy Optimization (Con-GRPO) then employs an Evidence-Aware Consistency Reward to reinforce reasoning that remains grounded in the image and logically supports the final answer. The resulting MedAD-R1 achieves state-of-the-art performance on MedAD-38K and consistently outperforms all evaluated baselines across five source-disjoint external datasets spanning diverse imaging modalities. Beyond accuracy, it achieves higher reasoning-answer consistency and visual-grounding scores across multiple backbones. With only 0.8B parameters, MedAD-R1 offers practical potential for resource-constrained deployment. These results demonstrate the generality of evidence-aware consistency optimization for interpretable MedAD. Project resources are available at https://github.com/zhtstar/MedAD-R1.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

    Jun 10, 2026Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci +6Medical Vision-Language ModelsRecent Vision-Language Models

  2. M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

    Jan 13, 2026Juntao Jiang, Jiangning Zhang, Yali Bi +7Medical ImagesDiagnostic Benchmark

  3. MedQA-MM: Shortcuts Behind Medical Visual Reasoning

    Sep 3, 2026Benlu Wang, Yifan Zhang, Jiaqing Yu +7Medical Visual Question AnsweringMultimodal Clinical Data