Large language models (LLMs) show promise in medical applications, but their ability to detect and correct errors in clinical texts remains under-evaluated, particularly beyond English. We introduce MedRECT, a bilingual benchmark for Japanese and English that formulates medical error handling as three subtasks: error detection, error sentence extraction, and error correction. MedRECT-ja contains 663 samples derived from the Japanese Medical Licensing Examinations, while the separately sourced MedRECT-en contains 458 samples curated from MEDEC. We evaluate 11 LLMs across 17 configurations that cover proprietary and open-weight models, medical-domain specialization, and multiple reasoning settings. Qwen3-32B scores higher in its thinking mode than in its non-thinking mode on error detection F1 and sentence extraction accuracy in both subsets, with sentence extraction accuracy higher by 24.5 percentage points on MedRECT-ja and 10.3 on MedRECT-en. Several leading general-purpose reasoning models outperform all three evaluated medical-domain models on these two subtasks. Most models have lower point estimates on the Japanese subset, although absolute scores are not directly comparable because the subsets differ in source material and error distributions. LoRA fine-tuning yields higher sentence extraction accuracy and higher point estimates on all three reference-based correction similarity metrics in both languages. MedRECT provides an open, reusable evaluation resource for studying medical error correction and reasoning across Japanese and English. Our dataset and code are available at https://github.com/pfnet-research/medrect.
Figures & tables
Figure 1: Overview of the MedRECT benchmark, illustrated with a MedRECT -ja sample (translated to English). Given a clinical text, models must (1) detect whether it contains a medical error, (2) identify the erroneous sentence, and (3) generate a corrected version. Here, the patient’s symptoms (difficulty distinguishing colors) and family history suggest congenital color vision deficiency, but a binocular vision test was ordered instead of the appropriate color vision test.
Figure 2: Examples from the MedRECT dataset showing different error types. Examples 1 and 2 show MedRECT -ja samples (translated to English for readability), while Example 3 shows a native MedRECT -en sample derived from MEDEC ( Ben Abacha et al., 2025 ) . Each example highlights the erroneous sentence (colored background), the key error in red , and the reference correction in green .
MedRECT -ja
MedRECT -en
Total samples
663
458
With errors
367 (55.4%)
243 (53.1%)
Without errors
296 (44.6%)
215 (46.9%)
Error Type Distribution
Diagnosis
77 (21.0%)
98 (40.3%)
Monitoring/management
79 (21.5%)
17 (7.0%)
Table 1: Dataset statistics for MedRECT -ja and MedRECT -en
MedRECT -ja
MedRECT -en
Model
Error Detection
Sentence Extraction
Reference-based correction similarity
Error Detection
Sentence Extraction
Reference-based correction similarity
F1
Acc. (%)
ROUGE-1
BERTScore
BLEURT
Avg.
F1
Acc. (%)
ROUGE-1
BERTScore
BLEURT
Avg.
Reasoning models
GPT-5
.758
83.7
.502
.712
.518
.577
.818
96.3
.689
.848
.697
.745
o3
.764
71.4
.454
.636
.458
.516
.852
87.7
.631
.781
.644
.685
Claude Sonnet 4
.795
82.3
.559
.755
.547
.620
.784
84.0
.642
.800
.643
.695
Table 2: Bilingual performance for 17 configurations of 11 LLMs on MedRECT -ja and MedRECT -en. Parentheses indicate reasoning effort levels for gpt-oss (high, medium, or low) and reasoning modes for Qwen3-32B (think or no-think).
Diagnosis
Monitoring/
Physical
Procedures/
Test
Medication
History
Medication
management
findings
intervention
interpretation
selection
taking
dosage
ja
en
ja
en
ja
en
ja
en
ja
en
ja
en
ja
en
ja
en
#samples
77
98
79
17
72
2
40
38
37
12
30
70
22
1
8
3
Reasoning models
GPT-5
90.9
99.0
89.9
100.0
72.2
100.0
97.5
89.5
70.3
100.0
100.0
94.3
45.5
100.0
100.0
100.0
o3
80.5
94.9
73.4
58.8
56.9
100.0
90.0
76.3
56.8
91.7
100.0
88.6
27.3
100.0
100.0
100.0
Table 3: Error sentence extraction accuracy (%) by error type on MedRECT -ja and MedRECT -en. For gpt-oss, only medium reasoning effort is shown. The four samples categorized as Other are excluded. Results based on fewer than 10 samples should be interpreted cautiously. Best values are bolded only for columns with at least 10 samples.
82-year-old female admitted to palliative care ward due to pancreatic cancer with liver metastasis. Progressive appetite loss and decreased food intake for 1 week […]. Physical findings: decreased skin turgor, poor oral hygiene and dryness […]. Blood findings: […] Hb 9.2g/dL […]. Biochemistry: […] albumin 2.8g/dL […]. Nutrition support team proposed gastrostomy for nutritional support. Plan to explain to attending physician and obtain consent.
80-year-old female diagnosed with lung cancer and multiple bone metastases 1 month ago. She chose not to receive active treatment and receives home visits […]. She mainly stays in bed during the day. Living alone since husband’s death at age 65, with distant contact from son […]. She told the visiting physician "I want to end it all." The physician asked "What is the most distressing physical symptom you have right now?" The patient did not respond and remained lying in bed.
45-year-old woman presents complaining of weakness that has progressively worsened over the past 2 weeks. She states that she has a hard time lifting both her arms but that they function normally […]. On examination, patient has 2/5 muscle strength on shoulder shrug and arm abduction bilaterally, but all other neurological exam findings are normal. You notice some skin changes […]. She states that she has had a rash around her eyes as well as on her lower face, going down to her neck and chest. She notes that the rashes started around the same time as the weakness began. Patient is diagnosed with dermatomyositis after labs are drawn and a complete blood count and basic metabolic panel are normal.
Reference Output
Nutrition support team should prioritize consideration of non-invasive nutrition management methods.
The physician responded "You’re feeling distressed. Could you tell me more about your feelings?"
CORRECT
GPT-5
Nutrition support team judged that gastrostomy is inappropriate given the terminal cancer with limited prognosis, and proposed non-invasive nutritional support prioritizing symptom palliation with oral intake strategies, oral care, and supplementary fluids as needed.
The physician first assessed for suicidal ideation and urgency (specific plans, means, intent), ensured safety, and then comprehensively assessed physical, psychological, social, and spiritual distress.
Patient is diagnosed with dermatomyositis after labs show elevated muscle enzymes (e.g., creatine kinase), while the complete blood count and basic metabolic panel may be normal.
Claude Sonnet 4
Nutrition support team proposed oral intake support for symptom palliation.
The physician asked "Could you tell me more about those feelings?"
Patient is diagnosed with dermatomyositis after labs are drawn including elevated muscle enzymes and positive autoantibodies , while complete blood count and basic metabolic panel are normal.
Qwen3-32B + LoRA (think)
Nutrition support team proposed coordination with home care , prioritizing continuation of oral intake above all else.
The physician responded "You have been carrying this alone and must be very tired."
CORRECT
Table 4: Model outputs on three selected MedRECT samples
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Data construction pipeline for MedRECT . MedRECT -ja (top) transforms JMLE questions through automated synthesis, consensus-based difficulty filtering, model deduplication, and LLM-as-a-Judge quality screening. MedRECT -en (bottom) applies the same screening framework with source-adapted criteria to the existing MEDEC MS test subset. Red numbers indicate samples removed at each step.
MedRECT -ja
MedRECT -en
Original samples
720
597
Retained samples
663 (92.1%)
458 (76.7%)
Exclusion Reasons
Ambiguous error
3
98
Extra elements
9
–
Multiple errors
21
24
Appendix
Table 5: Quality screening results and exclusion reasons. Reasons are non-mutually exclusive, so their counts can exceed the number of excluded samples.
Configuration change
Benchmark
Error detection F1 difference
Sentence extraction accuracy difference
Qwen3-32B: (no-think) → (think)
MedRECT -ja
+0.086 [0.051, 0.124]
+24.5 [17.5, 31.7]
Qwen3-32B: (no-think) → (think)
MedRECT -en
+0.037 [0.008, 0.066]
+10.3 [4.9, 15.6]
Qwen3-32B (think): without LoRA → with LoRA
MedRECT -ja
+0.020 [-0.011, 0.051]
+9.0 [2.8, 15.0]
Qwen3-32B (think): without LoRA → with LoRA
MedRECT -en
-0.012 [-0.040, 0.016]
+7.4 [2.5, 12.3]
Appendix
Table 6: Paired source-cluster bootstrap differences for focused within-model configuration changes. Each difference is the right-hand configuration minus the left-hand configuration. Brackets show 95% percentile confidence intervals from 20,000 resamples. Sentence extraction differences are percentage points.
Model
Error Detection
Sentence Extraction
Reference-based correction similarity †
F1
Acc. (%)
Acc. (%)
ROUGE-1
BERTScore
BLEURT
Avg.
MEDEC Paper Results
Medical Doctor #1
-
81.3
76.7
0.420
0.513
0.539
0.491
Medical Doctor #2
-
68.9
64.6
0.685
0.698
0.650
0.678
Reasoning models
GPT-5
0.780
71.7
90.7
0.655
0.672
0.671
0.666
Appendix
Table 7: Performance on the unfiltered MEDEC MS test subset. Parentheses give gpt-oss reasoning effort or Qwen3-32B reasoning mode.
As Large Language Models (LLMs) are increasingly deployed in healthcare settings, accurate error detection and correction in generated or existing text becomes critical, as even minor mistakes can pose risks to patient safety. Existing methods for error detection and correction, including automated checks and heuristic-based approaches, do not generalize well across unseen datasets. In this paper, we propose MedGuards as a medical safety guardrail, which is a new framework that treats medical error detection and correction as a multi-agent in-context learning task. Specialized agents separately detect, localize, and correct errors, while a confidence-guided arbitration mechanism resolves disagreements using reasoning traces and confidence scores. This design enhances interpretability, robustness, and adaptability, without requiring additional training of the base LLMs. Additionally, we introduce the Keyword-Prioritized Correction Score (KPCS), a new evaluation metric that considers whether critical keywords within the reference text are generated correctly, providing a more comprehensive assessment than conventional metrics. Experiments across four multilingual medical datasets consisting of clinical notes demonstrate significant improvements by the proposed framework across several metrics and models. Our aim is to enable safer deployment of LLMs in real-world healthcare applications. For reproducibility, we make our code publicly available at https://github.com/congboma/MedGuards.
Congbo Ma, Hu Wang, Yichun Zhang +1
New York University Abu Dhabi · Khalifa University · New York University
Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of medicine. We argue that robust cross-lingual medical transfer requires Hindi reasoning. To this end, we introduce HiMed, a Hindi reasoning medical corpus and benchmark suite covering both Western and Indian medicine. We further propose HiMed-8B, a Hindi-form medical reasoning LLM, through the design of decaying scaffolding reward. Extensive experiments demonstrate improvement in Hindi medical reasoning performance and reduction in the English--Hindi accuracy gap. Ablation studies validate the contribution of each training stage and reward component. All data and code are available on GitHub: https://github.com/FreedomIntelligence/HiMed.
Dingfeng Jiang, Han Yan, Chenze Ma +12
1The Chinese University of Hong Kong, Shenzhen · 2Indian Institute of Technology (Banaras Hindu University) Varanasi · 3Tongji University +4
Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general domains such as mathematics, failing to address medical reasoning -- which is uniquely characterized by safety criticality, knowledge intensity, and diverse error patterns. Without a reliable medical PRM evaluation framework, we cannot quantify models' error detection capabilities in clinical reasoning, leaving their safety in real-world healthcare applications unverified. We propose MedPRMBench, the first process-level reward model benchmark for the medical domain. Built through a three-phase pipeline based on Clinical Reasoning Blueprints (CRBs), MedPRMBench systematically generates high-quality evaluation data from seven medical QA sources, covering 14 fine-grained error types across three categories (Simplicity, Soundness, and Sensitivity) with the first 4-level severity grading system to quantify clinical impact. The benchmark comprises 6{,}500 questions with 13{,}000 reasoning chains and 113{,}910 step-level labels, plus 6{,}879 questions for training. Our medical PRM baseline achieves an 87.1% overall PRMScore -- substantially surpassing all baselines -- and serves as a plug-and-play verifier that improves downstream medical QA accuracy by 3.2--6.7 percentage points. Systematic evaluation spanning proprietary frontier models, open-source reasoning models, and medical-specialized models reveals critical weaknesses in current models' medical reasoning error detection capabilities, providing clear directions for future PRM improvement.