cs.CLNov 1, 2025

MedRECT: A Bilingual Medical Reasoning Benchmark for Error Correction in Clinical Texts

Authors: Naoto Iwase, Hiroki Okuyama, Junichiro Iwasawa

Organizations: Preferred Networks, Inc., Tokyo, Japan · School of Medicine, Nagoya University, Nagoya, Japan

Abstract

Large language models (LLMs) show promise in medical applications, but their ability to detect and correct errors in clinical texts remains under-evaluated, particularly beyond English. We introduce MedRECT, a bilingual benchmark for Japanese and English that formulates medical error handling as three subtasks: error detection, error sentence extraction, and error correction. MedRECT-ja contains 663 samples derived from the Japanese Medical Licensing Examinations, while the separately sourced MedRECT-en contains 458 samples curated from MEDEC. We evaluate 11 LLMs across 17 configurations that cover proprietary and open-weight models, medical-domain specialization, and multiple reasoning settings. Qwen3-32B scores higher in its thinking mode than in its non-thinking mode on error detection F1 and sentence extraction accuracy in both subsets, with sentence extraction accuracy higher by 24.5 percentage points on MedRECT-ja and 10.3 on MedRECT-en. Several leading general-purpose reasoning models outperform all three evaluated medical-domain models on these two subtasks. Most models have lower point estimates on the Japanese subset, although absolute scores are not directly comparable because the subsets differ in source material and error distributions. LoRA fine-tuning yields higher sentence extraction accuracy and higher point estimates on all three reference-based correction similarity metrics in both languages. MedRECT provides an open, reusable evaluation resource for studying medical error correction and reasoning across Japanese and English. Our dataset and code are available at https://github.com/pfnet-research/medrect.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction

    Jun 24, 2026Congbo Ma, Hu Wang, Yichun Zhang +1Correction

  2. HiMed: Incentivizing Hindi Reasoning in Medical LLMs

    May 23, 2026Dingfeng Jiang, Han Yan, Chenze Ma +12Multilingual Healthcare Q&AClinical Reasoning Training