HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
Organizations: The Chinese University of Hong Kong, Shenzhen · Shenzhen Research Institute of Big Data · Shenzhen Loop Area Institute · National Health Data Institute, Shenzhen
Abstract
Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.
Figures & tables
| Group | Method | External (Teacher) Guidance | Cold Start | Beyond Imitation | Late Imitation Dependence | Output Diversity |
| w/ SFT | SFT [ 1 , 2 , 3 , 4 ] | teacher outputs (SFT) | easy | ✗ | high | low |
| SFT+RL [ 7 , 10 ] | teacher outputs (SFT) | easy | ✓ | low | medium | |
| RL-only | Pure On-Policy RL [ 22 ] | none | hard | ✓ | low | high |
| Standard Mixed-Policy RL [ 21 ] | persistent teacher outputs | easy | ✓ | high | medium | |
| OnePO (Ours) | transient teacher outputs | easy | ✓ | low | high |
| Open-ended | Closed-ended | |||||
| Model | HealthBench Professional | HealthBench (Total) | HealthBench (Hard) | Medbullets (5-option) | MMLU-Pro (Med) | MedXpertQA (Text) |
| Representative Baselines | ||||||
| GPT-6 Astra (high) | 62.5 | 57.8 | 32.1 | 88.9 | 87.6 | 64.4 |
| GPT-5.2 (high) | 52.8 | 59.9 | 40.5 | 88.9 | 90.0 | 54.2 |
| Gemini 3.8 Flash | 49.9 | 54.3 | 26.4 | 91.9 | 88.4 | 63.0 |
| DeepSeek-V4.1-Flash (high) | 48.5 | 57.7 | 32.7 | 86.0 | 85.8 | 51.6 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Open-ended | Closed-ended | ||||
| Setting | HealthBench (Total) | HealthBench (Hard) | Medbullets | MMLU-Pro | MedXpertQA |
| Qwen3-8B-Base | 22.1 | 0.0 | 30.7 | 41.2 | 11.7 |
| w/ Verifiable Only | 24.2 | 3.2 | 64.5 | 80.5 | 24.8 |
| w/ Rubric Only | 63.7 | 36.6 | 44.8 | 67.0 | 12.5 |
| w/ Mixed (Ours) | 65.4 | 39.1 | 64.0 | 81.2 | 24.9 |
| Open-ended (HealthBench) | Closed-ended | |||||
| Teacher Source | Setting | Total | Hard | Medbullets | MMLU-Pro(Med) | MedXpertQA |
| None | Pure RL | 59.8 | 25.2 | 48.1 | 75.0 | 20.0 |
| GPT-5 Chat | SFT | 30.4 | 4.8 | 49.7 | 72.5 | 21.7 |
| SFT+RL | 63.6 | 37.3 | 61.9 | 78.2 | 21.5 | |
| OnePO | 65.4 | 39.1 | 64.0 | 81.2 | 24.9 | |
| DeepSeek-V3.2 (thinking) | SFT | 36.8 | 2.4 | 56.6 | 71.0 | 17.2 |
| Floor | HealthBench | Medbullets | MMLU-Pro (Med) |
| 0.00 | 49.1 | 49.7 | 73.1 |
| 0.01 | 58.3 | 59.7 | 77.2 |
| 0.05 | 64.8 | 63.2 | 79.5 |
| 0.10 | 65.4 | 64.0 | 81.2 |
| 0.20 | 62.1 | 61.5 | 79.8 |
| 0.50 | 60.2 | 59.1 | 77.3 |
| Method | Steps to reward |
| OnePO | 7 |
| OnePO w/ KL penalty | 7 |
| OnePO w/ KL loss | 8 |
| OnePO w/ KL loss + KL penalty | 8 |
| OnePO w/o probability floor | Fail |
| OnePO w/o rescaling | Fail |
| Method | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 |
| OnePO | -6.9735 | -6.1488 | -4.2228 | -2.7612 | -1.0261 | -0.2838 | -0.0156 | -0.0122 |
| OnePO w/ KL loss + KL penalty | -6.9735 | -6.6096 | -5.5103 | -3.5550 | -2.2576 | -1.0866 | -0.3733 | -0.0406 |
| OnePO w/o rescaling | -6.9735 | -7.1113 | -7.1156 | -7.1270 | -6.9345 | -6.9741 | -6.5402 | -6.6415 |
| Method | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | Avg. |
| OnePO | 0.0000 | 0.0645 | 0.0608 | 0.0280 | 0.0038 | 0.0047 | 0.0072 | 0.0064 | 0.0219 |
| OnePO w/ KL loss + KL penalty | 0.0000 | 0.0410 | 0.0238 | 0.0072 | -0.0125 | 0.0015 | 0.0033 | 0.0018 | 0.0083 |
| OnePO w/o rescaling | 0.0000 | 0.0533 | 0.0590 | 0.0755 | 0.1160 | 0.1456 | 0.1471 | 0.1779 | 0.0968 |
| Method | Writing | Law | ||
| WritingBench | CreativeWriting-v3 | LexEval | LawBench | |
| Qwen3-4B-Base | 33.7 | 23.9 | 18.6 | 41.8 |
| w/ Pure RL | 67.4 | 41.1 | 46.5 | 60.8 |
| w/ SFT+RL | 71.2 | 45.6 | 50.7 | 64.5 |
| w/ OnePO | 74.6 | 47.8 | 51.9 | 67.4 |
| Metric vs. Human Label | 8B Grader (Training) | GPT-4.1 Grader (Evaluation) |
| F1 score | 0.81 | 0.85 |
| Cohen’s kappa | 0.62 | 0.72 |
| Model | 8B Grader | GPT-4.1 Official | GPT-5 Chat Grader |
| Qwen3-8B | 48.5 | 45.9 | 41.8 |
| Qwen3-8B w/ OnePO | 70.7 | 67.2 | 65.1 |
| Gain | +22.2 | +21.3 | +23.3 |
| Train step | 0 | 20 | 40 | 60 | 80 | 100 | 120 |
| Open-ended reward | 0.34 | 0.69 | 0.69 | 0.70 | 0.70 | 0.75 | 0.77 |
| Closed-ended reward | 0.29 | 0.38 | 0.48 | 0.52 | 0.55 | 0.58 | 0.59 |