Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized language models that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare foundation model trained through recursive self-improvement (RSI). In each RSI round, the RSI harness runs the current model on healthcare benchmarks, evaluates the trajectories to locate capability gaps, and refines the training mixture by synthesizing training data. On 6 healthcare benchmarks, Cura 1T scores highest on MedAgentBench, HealthBench Professional, HealthBench Hard, MedXpertQA text, and AgentClinic, and second on MedXpertQA multimodal. It preserves performances on out-of-domain reasoning and agentic benchmarks including AIME, GPQA-Diamond, and τ2-Bench.
Figures & tables
Benchmark
Base
Cura 1T
Δ
MedAgentBench
0.847
0.940
+0.093
HealthBench Professional
0.503
0.662
+0.159
HealthBench Hard
0.222
0.368
+0.146
MedXpertQA
0.569
0.655
+0.086
AgentClinic
0.754
0.796
+0.042
Table 1: Improvement of Cura 1T over its base model. Metrics differ by benchmark (rubric score for HealthBench, exact-letter pass@1 for MedXpertQA, task success for AgentClinic and MedAgentBench).
Figure 1: Left: recursive self-improvement loop for Cura 1T. A reviewer approves the plan before training and makes the keep, revert, or deploy decision after evaluation. Right: data refinement pipeline and training stack.
Figure 2: Cura 1T against frontier models and the Kimi-K2.6 base across six healthcare benchmarks: MedAgentBench ( Jiang et al., 2025 ) , HealthBench Professional and Hard ( Arora et al., 2025 ; OpenAI, 2026 ) , MedXpertQA text and multimodal ( Zuo et al., 2025 ) , and AgentClinic ( Schmidgall et al., 2024 ) .
Model
Intervention
Overall
Decision
Kimi-K2.6
−
0.847
−
Round 1
Behavior calibration
0.943
keep
Round 2
Behavior calibration + retention
0.967
keep
Round 3
Harness bug fix
0.973
keep
Cura 1T
Consolidated data mixture
0.940
release
Claude Opus 4.8
−
0.937
−
Table 2: MedAgentBench task success. Bold marks the best score and underline the second-best distinct score.
Model
Intervention
Professional
Hard
Decision
Kimi-K2.6
−
0.503
0.222
−
Round 1
Behavior calibration
0.601
0.332
revert
Round 2
Behavior calibration, cleaned
0.634
0.372
keep
Cura 1T
Consolidated data mixture
0.662
0.368
release
Table 3: HealthBench rubric scores at T=1.0 . Bold marks the best score and underline the second-best distinct score in each column.
Model
Intervention
Text
Multimodal
Overall
Decision
Kimi-K2.6
−
0.484
0.672
0.569
−
Round 1
Reasoning correction
0.447
0.656
0.541
revert
Round 2
Reasoning correction, extended
0.454
0.657
0.545
revert
Round 3
Knowledge injection + retention
0.521
0.703
0.603
keep
Round 4
Data mixture curation
0.560
0.728
0.636
keep
Round 5
Data mixture curation, capped
−
−
0.440
revert
Table 4: MedXpertQA split and overall pass@1 at T=1.0 . The overall score weights 2,450 text questions and 2,000 multimodal questions. Bold marks the best score and underline the second-best distinct score in each column.
Model
Intervention
MedQA
MedQA Ext
NEJM
NEJM Ext
Overall
Decision
Kimi-K2.6
−
0.869
0.827
0.400
0.567
0.754
−
Round 2
Behavior calibration + retention
0.841
0.841
0.800
0.717
0.807
keep
Cura 1T
Consolidated data mixture
0.879
0.850
0.800
0.625
0.796
release
Claude Opus 4.8
−
0.841
0.874
0.800
0.608
0.794
−
GPT-5.5
−
0.832
0.808
0.467
0.358
0.684
−
Table 5: AgentClinic pass@1 by subset under the tool-native protocol. Bold marks the best score and underline the second-best distinct score in each column.
Figure 3: Improvement map from the base model to Cura 1T. Values are changes from benchmark-specific bases; solid and dashed red arrows mark retained and reverted interventions.
Figure 4: Out-of-domain evaluation results for Cura 1T.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Role in the loop
Output
Plan
Define target behavior, metrics, data recipe, and candidate hyperparameters.
Training specifications
Train
Run SFT to verify the mixture and hyperparameters, RL to improve the policy against reward signals, and SDFT to consolidate the final model.
Candidate model
Evaluate
Run benchmark harnesses and collect graded trajectories and failure summaries.
Failed trajectories and metrics
Refine
Categorize failures, synthesize targeted data, curate the next mixture, validate candidate rows, and suggest next steps.
Validated data refinement
Reasoning Correction
Use a corrected reasoning pattern to synthesize new cases in different clinical contexts.
Reasoning rows
Knowledge Injection
Use verified knowledge to synthesize new questions that require the missing concept.
Source-grounded knowledge rows
Appendix
Table 6: Components of the recursive self-improvement loop in Figure 1 . The table summarizes how evaluation records are converted into refined data mixtures and how SFT, RL, and SDFT divide the training step.