InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision
Organizations: The Hong Kong Polytechnic University · PolyU-Daya Bay Technology and Innovation Research Institute
Abstract
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.
Figures & tables
| Model | MMMU Med. | MMMU-Pro Med.-10 | MedXQA | PMC-VQA clean | OmniMed VQA | Avg. |
|---|---|---|---|---|---|---|
| General-purpose models | ||||||
| Gemini-3-Pro | 81.84 | 74.13 | 74.85 | 70.40 | 84.55 | 77.15 |
| Gemini-3-Flash | 81.51 | 70.63 | 68.00 | 69.75 | 84.45 | 74.87 |
| Claude-Opus-4.7 | 82.47 | 70.63 | 69.40 | 67.15 | 81.60 | 74.25 |
| GPT-5 | 83.39 | 70.90 | 71.70 | 67.30 | 76.40 | 73.94 |
| GLM-5V-Turbo | 73.91 | 64.69 | 53.25 | 66.55 | 81.95 | 68.07 |
| Parameters | Vision-encoder adaptation | Warmup– stable CPT | Learning-rate decay | Post- training | MMMU Med. | MMMU-Pro Med.-10 | MedXQA | PMC-VQA clean | OmniMedVQA | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|
| 4B | – | – | – | SFT | 69.53 | 56.64 | 39.35 | 62.15 | 92.12 | 63.96 |
| 1B | 30B | 2.5B | SFT | 71.18 | 56.99 | 42.70 | 63.00 | 90.11 | 64.80 | |
| 1B | 55B | 2.5B | SFT | 70.55 | 55.94 | 43.05 | 63.90 | 90.04 | 64.69 | |
| 1B | 30B | 2.5B | SFT+RLVR | 72.43 | 58.39 | 46.25 | 65.80 | 90.79 | 66.73 | |
| 27B | – | – | – | SFT | 75.51 | 68.88 | 56.55 | 66.20 | 92.44 | 71.92 |
| 1B | 8B | 2.5B | SFT | 78.88 | 66.78 | 57.45 | 66.20 | 92.69 | 72.40 |
| SFT target | MMMU Med. | MMMU-Pro Med.-10 | MedXQA | PMC-VQA clean | OmniMedVQA | Avg. |
|---|---|---|---|---|---|---|
| Answer only | 58.56 | 40.21 | 31.25 | 65.70 | 83.19 | 55.78 |
| Random generated response | 68.04 | 52.10 | 34.85 | 59.25 | 75.60 | 57.97 |
| Stability-aware selected response | 69.81 | 51.75 | 35.00 | 59.20 | 77.00 | 58.55 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Tokens | Share |
|---|---|---|
| Medical e-books | 1.000B | 40.00% |
| Biomedical scientific literature | 0.500B | 20.00% |
| General multimodal replay | 0.396B | 15.84% |
| Medical web articles | 0.094B | 3.76% |
| Medical web exercises and selected VQA | 0.510B | 20.40% |
| Total | 2.500B | 100.00% |
| Dimension | Score | Criterion | Description |
|---|---|---|---|
| Image–question alignment | 0 | Non-medical image | The image is unrelated to the medical context required by the question. |
| 0 | Modality mismatch | The visual modality required by the question is inconsistent with the provided image, e.g., a question referring to CT while the input is an X-ray. | |
| 0 | Anatomical mismatch | The organ, anatomical region, or laterality specified by the question is absent from the provided image(s). | |
| 0 | Missing image | The question refers to an image or panel that is not provided, or the available images are insufficient for answering the question. | |
| 1 | Image-independent question | The question can be answered without using the visual input. | |
| Information completeness | 0 | Invalid answer space | The candidate answers do not contain a valid option or are not mutually compatible with the intended question. |
| Category | Examples | Tokens | Share | Vision | Text |
|---|---|---|---|---|---|
| General multimodal instruction | 50,000 | 26.021M | 10.76% | 16.909M | 9.112M |
| Medical visual question answering | 90,734 | 52.644M | 21.77% | 41.699M | 10.945M |
| Medical reasoning data | 49,808 | 54.868M | 22.69% | 16.469M | 38.399M |
| Stability-aware supervision | 106,183 | 108.316M | 44.79% | 45.382M | 62.933M |
| Total | 296,725 | 241.848M | 100.00% | 120.459M | 121.389M |
| Stage | Count | Retention |
|---|---|---|
| Input questions | 176,948 | 100.00% |
| Generated candidates ( per question) | 1,415,584 | 100.00% |
| Candidates passing answer checking | 704,046 | 49.74% |
| Questions with at least one valid candidate | 121,544 | 68.69% |
| Long-form examples exported after stability filtering | 108,293 | 61.20% |
| Final stability-aware SFT examples | 106,183 | 60.01% |
| Stage | Trainable components | Epochs | Max. length | Batch | Learning-rate schedule |
|---|---|---|---|---|---|
| Vision-encoder adaptation | Vision encoder, projector | 1 | 8,192 | 256 / 2 | 3% warmup to , cosine decay to |
| Warmup–stable CPT | All components | 1 | 12,288 | 256 / 1 | 3% warmup to , then constant |
| Learning-rate decay | All components | 1 | 12,288 | 256 / 1 | Cosine decay from to |
| Supervised fine-tuning | All components | 5 | 9,000 | 64 / 2 | 10% warmup, then cosine decay to zero. Peak for language model and projector and for vision encoder |
| Hyperparameter | Value |
|---|---|
| Learning rate | |
| Global batch size | 128 |
| Training epochs | 3 |
| Rollouts per prompt | 8 |
| Sampling temperature | 1 |
| Format-to-accuracy reward ratio |
| Category | Source | Train | Test | Total |
|---|---|---|---|---|
| Medical VQA | PMC-VQA | 7,210 | 790 | 8,000 |
| MedSynVQA | 5,371 | 629 | 6,000 | |
| GMAI-VL-5.5M | 2,721 | 279 | 3,000 | |
| Internally curated medical data | 1,848 | 206 | 2,054 | |
| OmniMedVQA + RadImageNet | 426 | 49 | 475 | |
| Text QA | MedQA (English USMLE) | 1,800 | 200 | 2,000 |