SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis
Organizations: East China Normal University, Shanghai, China · Alibaba Group, Hangzhou, China
Abstract
Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.
Figures & tables
| Model | Method/Dataset | Scale | AVG | MedMCQA | MedQA | PubMedQA | MMLU-Pro | GPQA |
|---|---|---|---|---|---|---|---|---|
| (med) | (med) | |||||||
| Qwen3-4B-Base | Few-shot Prompting | NA | 26.15 | 29.14 | 29.69 | 35.20 | 14.66 | 22.05 |
| Qwen3-4B NoThinking | NA | NA | 62.54 | 56.78 | 63.63 | 71.00 | 64.63 | 56.67 |
| Qwen3-4B Thinking | NA | NA | 66.18 | 59.81 | 72.03 | 72.60 | 71.60 | 54.87 |
| Open Source Datasets | ||||||||
| Qwen3-4B-Base | m1k ( 2025a ) | 1k | 63.40 | 59.26 | 72.19 | 74.20 | 67.49 | 43.85 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Strategy | MedNLI | MTS |
|---|---|---|
| Wrap | 72.08 | 83.62 |
| CoT-Self-Instruct | 84.60 | 83.16 |
| SAGE | 87.27 | 85.02 |
| Model | MedMCQA | MedQA | PubMedQA | MMLU-Pro | GPQA | AVG |
|---|---|---|---|---|---|---|
| (med) | (med) | |||||
| SAGE-4B | 64.98 | 70.07 | 80.80 | 63.19 | 61.54 | 68.12 |
| 8B Scale | ||||||
| UltraMedical-8B Zhang et al. (2024) | 59.17 | 71.64 | 70.30 | 61.43 | 50.26 | 62.56 |
| m1-7B Huang et al. (2025a) | 61.49 | 71.96 | 74.00 | 63.65 | 46.67 | 63.55 |
| MedReason-8B Wu et al. (2025) | 60.17 | 70.78 | 78.50 | 65.08 | 52.56 | 65.42 |
| Strategy | TruthfulQA | MedHALT |
|---|---|---|
| Wrap | 32.66 | 90.84 |
| CoT-SI | 22.61 | 91.59 |
| SAGE | 20.60 | 89.28 |
| Strategy | Data Source | Tokens | MedMCQA | MedQA | PubMedQA | MMLU-Pro | GPQA | AVG |
|---|---|---|---|---|---|---|---|---|
| Genie | Scope Notes | 1.69M | 49.32 | 51.77 | 63.90 | 48.99 | 35.90 | 49.98 |
| Genie | Notes + Distilled Q&A | 3.51M | 46.86 | 55.15 | 75.30 | 52.31 | 35.13 | 52.95 |
| Genie | PubMed Abstracts | 3.51M | 52.38 | 48.63 | 61.00 | 51.21 | 40.00 | 50.64 |
| SAGE | Notes + Distilled Q&A | 3.51M | 64.98 | 70.07 | 80.80 | 63.19 | 61.54 | 68.12 |
| Extraction Method | Precision | Recall |
|---|---|---|
| Naive String Matching | 36.52 | 55.30 |
| SAGE Mapping (Ours) | 55.68 | 12.21 |
| Extraction Method | Avg Time / Sample | Privacy | MedMCQA | MedQA | PubMedQA | MMLU-Pro | GPQA | AVG |
|---|---|---|---|---|---|---|---|---|
| Naive Matching | 1.15s | Offline | 63.81 | 69.60 | 70.40 | 61.95 | 54.36 | 64.02 |
| MeSH on Demand | 48.24s | Online API | 66.08 | 67.16 | 83.20 | 62.74 | 61.54 | 68.14 |
| SAGE Mapping | 1.31s | Offline | 64.98 | 70.07 | 80.80 | 63.19 | 61.54 | 68.12 |
| Filtering Strategy | MedMCQA | MedQA | PubMedQA | MMLU-Pro | GPQA | AVG |
|---|---|---|---|---|---|---|
| Contradiction Filtering | 64.67 | 68.03 | 70.80 | 60.72 | 58.46 | 64.54 |
| SAGE (Consistency + Atomic Gen.) | 64.98 | 70.07 | 80.80 | 63.19 | 61.54 | 68.12 |
| Model | Method/Dataset | Scale | AVG-5 | AVG-10 | Lancet | MedBul_op4 | MedBul_op5 | MedXpert | NEJM |
| Qwen3-4B-Base | Few-shot Prompting | NA | 26.15 | 24.88 | 30.58 | 29.55 | 24.03 | 11.18 | 22.72 |
| Qwen3-4B NoThinking | NA | NA | 62.54 | 54.15 | 59.47 | 50.32 | 48.05 | 13.04 | 57.88 |
| Qwen3-4B Thinking | NA | NA | 66.18 | 57.98 | 62.86 | 59.09 | 50.97 | 13.73 | 62.19 |
| Open Source Datasets | |||||||||
| Qwen3-4B-Base | m1k | 1k | 63.40 | 56.98 | 61.65 | 59.74 | 54.87 | 17.05 | 59.54 |
| Qwen3-4B-Base | m23k | 23k | 67.39 | 61.44 | 65.05 | 67.53 | 61.04 | 20.50 | 63.35 |
| Method | AVG-5 | MedMCQA | MedQA | PubMedQA | MMLU-Pro (med) | GPQA (med) |
|---|---|---|---|---|---|---|
| SAGE (thinking) | ||||||
| SAGE (non-thinking) | ||||||
| CoT-Self-Instruct (thinking) | ||||||
| CoT-Self-Instruct (non-thinking) |
| Method | AVG-5 | MedMCQA | MedQA | PubMedQA | MMLU-Pro (med) | GPQA (med) |
|---|---|---|---|---|---|---|
| SAGE | ||||||
| CSI | ||||||
| WRAP |
| Dataset | Train-Set | SAGE |
|---|---|---|
| MedMCQA | ||
| MedQA | ||
| PubMedQA |