Rufus-Air: An Open LLM Post-Training Recipe
Organizations: Amazon
Abstract
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.
Figures & tables
| Rufus -Air | GLM-4.5-Air | INTELLECT-3 | Nemotron-3 | Qwen3.5 | GPT-OSS | Ring-flash | Solar-Open | Sarvam | Mistral | |
| Total params | 106B | 106B | 106B | 120B | 122B | 117B | 100B | 102B | 105B | 119B |
| Active params | 12B | 12B | 12B | 12B | 10B | 5B | 6B | 12B | 10B | 7B |
| Instruction following & alignment | ||||||||||
| IFBench (prompt strict) | 76.9 | 33.6 | 29.3 | 68.6 | 76.1 | 69.0 | – | 57.7 | – | 48.0 |
| IFEval (prompt strict) | 95.4 | 83.0 | 79.5 | 91.3 | 93.4 | 88.9 | – | 88.0 | 84.8 | 84.0 |
| Multi-challenge | 65.8 | 36.0 | 36.5 | 50.7 | 61.5 | 45.3 | – | 40.5 | – | – |
| Checkpoint | GPQA | AIME25 | AIME26 | LCBv6 | IFEval | IFBench | Multi-ch. | MCP-A | Tau2-Re | TB2.1 | SWE-V | BrowseC | Seal-0 | HLE-V | AH(HP) | AH(CW) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GLM-4.5-Air | 73.9 | 84.2 | 86.5 | 59.6 | 83.0 | 33.6 | 36.0 | 35.9 | 80.7 | 24.7 | 50.6 | 22.7 | 33.3 | 20.2 | 55.0 | 60.3 |
| SFT (3799) | -5.7 68.2 -5.7 | +6.6 90.8 +6.6 | +3.5 90.0 +3.5 | – | – | – | – | – | – | – | – | – | – | – | – | – |
| + Reasoning RL | +5.3 73.5 +5.3 | -2.8 88.0 -2.8 | -2.6 87.4 -2.6 | 68.6 | – | – | – | – | – | – | – | – | – | – | – | – |
| + Coding RL | – | – | – | +7.3 75.9 +7.3 | 90.5 | 63.8 | 31.1 | – | – | – | – | – | – | – | – | – |
| + IF RL | – | – | – | – | +4.0 94.5 +4.0 | +14.0 77.8 +14.0 | +24.7 55.8 +24.7 | 35.0 | 74.0 | – | – | – | – | – | – | – |
| + General Agent | – | – | – | – | – | – | – | +7.8 42.8 +7.8 | +9.8 83.8 +9.8 | 38.8 | 65.6 | – | – | – | – | – |
| IFEval | IFBench | GPQA | AIME 25 | AIME 26 | |
| GLM-4.5-Air | 83.00 | 33.60 | 73.90 | 84.20 | 86.50 |
| Rufus-Air SFT (3799) | 88.33 | 57.75 | 68.18 | 90.83 | 90.00 |
| Prompts | Prompt tokens | |||||
| Task | Data source | Verifier | # | % | # | % |
| Math | HF crawl; synthesized | Math-Verify (canonical match) | 57,736 | 47.7 | 6.34M | 25.3 |
| Science | HF multi-science crawl | Fuzzy string matching | 43,199 | 35.7 | 4.81M | 19.2 |
| Puzzles | Enigmata; ReasoningGym | Generated Python checkers | 20,226 | 16.7 | 13.91M | 55.5 |
| Total | 121,161 | 100 | 25.07M | 100 | ||
| Checkpoint | GPQA | AIME 25 | AIME 26 |
|---|---|---|---|
| Rufus-Air SFT | 68.18 | 90.83 | 90.00 |
| Reasoning RL | 73.50 | 88.02 | 87.40 |
| Source | Modality | Problems | % of set |
|---|---|---|---|
| EvolveCoder (synthetic) | functional | 14,743 | 50.1% |
| Nemotron (competitive) | stdin/stdout | 9,393 | 31.9% |
| Dolci | mixed | 3,109 | 10.6% |
| ADR (algorithmic) | stdin/stdout | 2,160 | 7.3% |
| Total | 53.2% fn / 46.8% i/o | 29,405 | 100% |
| Checkpoint | Pass@1 | Pass@8 |
|---|---|---|
| Reasoning RL | 68.6 | 85.1 |
| Coding RL | 75.9 | 87.4 |
| Checkpoint | IFEval | IFBench | Multi-challenge | AdvancedIF | GPQA | AIME 25 |
|---|---|---|---|---|---|---|
| Coding RL | 90.5 | 63.8 | 31.1 | 45.5 | 67.9 | 87.6 |
| IF RL | 94.5 | 77.8 | 55.8 | 62.9 | 71.1 | 86.4 |
| Checkpoint | MCP-Atlas | Tau2-Retail | AIME 26 | IFEval | GPQA | LiveCodeBench |
|---|---|---|---|---|---|---|
| IF RL | 34.95 | 74.00 | 87.08 | 94.73 | 73.36 | 71.21 |
| General Agent | 42.75 | 83.80 | 87.60 | 96.16 | 76.39 | 73.36 |
| Checkpoint | Terminal-Bench 2.1 | SWE-bench Verified | AIME 26 | IFEval | LiveCodeBench |
|---|---|---|---|---|---|
| General Agent | 38.76 | 65.60 | 87.60 | 96.16 | 73.36 |
| Coding Agent | 40.17 | 67.80 | 87.40 | 95.68 | 75.64 |
| Dataset | Original # | Valid # | Avg. Chosen Score | Avg. Rejected Score | Avg. Chosen Tokens | Avg. Rejected Tokens |
|---|---|---|---|---|---|---|
| Arena Human Preference | 84,402 | 53,502 | 17.06 | 12.08 | 1055.1 | 623.7 |
| HelpSteer3 | 38,459 | 29,506 | 9.60 | 2.95 | 439.1 | 373.2 |
| HH-RLHF | 112,052 | 75,815 | -4.23 | -8.41 | 78.8 | 72.9 |
| Arena-Hard v2 (HP) | Arena-Hard v2 (CW) | |
| Search Agent | 83.06 | 38.56 |
| + RLHF | 89.05 | 52.97 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Version |
|---|---|
| Megatron-LM | 3714d81 |
| SGLang | 24c9100 |
| Strands Agents | 1.50.2 |
| PyTorch [ 74 ] | 2.9.1 (cu129) |
| TransformerEngine [ 67 ] | 2.10.0 |
| FlashAttention [ 21 ] | 2.7.4.post1 |
| Stage | Optimizer | LR | Batch | Samples/prompt | Opt. steps/rollout | Response budget | Nodes |
|---|---|---|---|---|---|---|---|
| SFT | AdamW | 4096 seq | n/a | n/a | 128K ctx | 64 | |
| Reasoning RL | AdamW | 256 prompts | 16 | 2 | 30K tok | 8 | |
| Coding RL | AdamW | 128 64 prompts | 64 | 8 | 64K 128K | 32 | |
| IF RL | AdamW | 256 prompts | 16 | 1 | 16K tok | 8 | |
| General Agent | AdamW | 48 prompts | 64 | 1 | 32K tok | 16 | |
| Coding Agent | AdamW | 32 prompts | 32 | 1 | 60K tok | 16 |
| Dataset | Terms | Released-data provenance |
|---|---|---|
| Ring-lite-sft-data [ 54 ] | Apache-2.0 | Aggregates open datasets including BigMath, DeepScaleR, and DAPO, whose problems originate from public competition and textbook sources. |
| hermes_reasoning_ tool_use [ 40 ] | Apache-2.0 | Derived from ShareGPT conversations originally collected from sharegpt.com . |
| ToolMind [ 111 ] | Apache-2.0 | Draws on xLAM, When2Call, glaive-function-calling-v2, ToolACE, and BUTTONInstruct; released responses were generated with DeepSeek-V2-Chat, Mixtral-8x22B-Instruct, and DeepSeek-V3. |
| tool-use-multiturn-reasoning [ 39 ] | Apache-2.0 | DeepSeek-R1 and QwQ-32B generated the released data. |
| ToolMind-Web-QA | Apache-2.0 | QA pairs follow rules derived from Wikipedia entity–relation graphs; search trajectories were generated by MiroThinker using Qwen3-235B-A22B-Thinking-2507. |
| Toucan-1.5M [ 108 ] | Apache-2.0 | Qwen3-32B, Kimi-K2, and GPT-OSS generated the released data. |
| Qwen3.5 | GPT-OSS | Ring-flash-2.0 | Solar-Open | Sarvam | Mistral-Small-4 | |
| 122B-A10B | 120B | 100B-A6B | 102B-A12B | 105B-A10B | 119B-A7B | |
| IFBench | [ 81 ] | [ 81 ] | – | [ 101 ] | – | [ 63 ] |
| IFEval | [ 81 ] | [ 81 ] | – | [ 73 ] | [ 86 ] | [ 123 ] |
| Multi-challenge | [ 81 ] | [ 87 ] | – | [ 101 ] | – | – |
| Arena-Hard v2 (HP) | [ 69 ] | [ 69 ] | – | – | – | – |
| AIME 25 | [ 69 ] | [ 71 ] | [ 53 ] | [ 73 ] | [ 86 ] | [ 63 ] |
| Dataset | Checkpoint | Iters | Search | Scrape | Python | Iters corr./wrong |
|---|---|---|---|---|---|---|
| BrowseComp | GLM-4.5-Air | 54.0 | 46.6 | 7.3 | 0.0 | 32.9 / 61.6 |
| Coding Agent | 70.6 | 47.6 | 22.1 | 0.9 | 42.8 / 85.1 | |
| Search Agent | 75.3 | 51.9 | 22.2 | 1.0 | 42.5 / 92.0 | |
| Seal-0 | GLM-4.5-Air | 18.0 | 11.3 | 6.4 | 0.3 | 17.9 / 18.0 |
| Coding Agent | 31.6 | 15.3 | 15.5 | 0.8 | 26.9 / 36.3 | |
| Search Agent | 29.8 | 13.8 | 14.9 | 1.0 | 25.0 / 34.4 |
| Domain | User simulator | Rufus -Air † | GLM-4.5-Air | INTELLECT-3 | Nemotron-3 |
|---|---|---|---|---|---|
| Sonnet 4 | 74.0 | 69.5 | 65.0 | 78.5 | |
| Sonnet 4.5 | 76.0 | 74.0 | 70.5 | 75.5 | |
| Airline | Sonnet 5 | 80.5 | 77.0 | 76.0 | 86.0 |
| Sonnet 4 | 69.1 | 57.0 | 68.4 | 68.2 | |
| Sonnet 4.5 | 84.6 | 74.8 | 73.7 | 82.2 | |
| Retail | Sonnet 5 | 86.6 | 80.7 | 75.7 | 86.4 |