Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.
Figures & tables
Figure 1 : Recipe overview. The post-training pipeline runs from the GLM-4.5-Air-Base checkpoint through eight stages to the final Rufus-Air model, each stage trained on the checkpoint the previous one produced.
Rufus -Air
GLM-4.5-Air
INTELLECT-3
Nemotron-3
Qwen3.5
GPT-OSS
Ring-flash
Solar-Open
Sarvam
Mistral
Total params
106B
106B
106B
120B
122B
117B
100B
102B
105B
119B
Active params
12B
12B
12B
12B
10B
5B
6B
12B
10B
7B
Instruction following & alignment
IFBench (prompt strict)
76.9
33.6
29.3
68.6
76.1
69.0
–
57.7
–
48.0
IFEval (prompt strict)
95.4
83.0
79.5
91.3
93.4
88.9
–
88.0
84.8
84.0
Multi-challenge
65.8
36.0
36.5
50.7
61.5
45.3
–
40.5
–
–
Table 1 : Performance comparison. All values are pass@1 [ 18 ] except the two Arena-Hard v2 rows, which are win rates; § 5.1 describes and cites every benchmark. The four leftmost models we evaluated ourselves under the same harness, and we bold the best of those four in each row; the other six models are from other public sources, listed in Appendix D . A dash marks a benchmark we did not run (first four columns) or that the source did not report (other six).
Checkpoint
GPQA
AIME25
AIME26
LCBv6
IFEval
IFBench
Multi-ch.
MCP-A
Tau2-Re
TB2.1
SWE-V
BrowseC
Seal-0
HLE-V
AH(HP)
AH(CW)
GLM-4.5-Air
73.9
84.2
86.5
59.6
83.0
33.6
36.0
35.9
80.7
24.7
50.6
22.7
33.3
20.2
55.0
60.3
SFT (3799)
-5.7 68.2 -5.7
+6.6 90.8 +6.6
+3.5 90.0 +3.5
–
–
–
–
–
–
–
–
–
–
–
–
–
+ Reasoning RL
+5.3 73.5 +5.3
-2.8 88.0 -2.8
-2.6 87.4 -2.6
68.6
–
–
–
–
–
–
–
–
–
–
–
–
+ Coding RL
–
–
–
+7.3 75.9 +7.3
90.5
63.8
31.1
–
–
–
–
–
–
–
–
–
+ IF RL
–
–
–
–
+4.0 94.5 +4.0
+14.0 77.8 +14.0
+24.7 55.8 +24.7
35.0
74.0
–
–
–
–
–
–
–
+ General Agent
–
–
–
–
–
–
–
+7.8 42.8 +7.8
+9.8 83.8 +9.8
38.8
65.6
–
–
–
–
–
Table 2 : Stagewise progression of Rufus-Air , assembled from the per-stage result tables. One row per checkpoint, in pipeline order: each row carries the benchmarks targeted by the stage that produced it (with the change against the row above) and, in the next stage’s target columns, the starting values that stage’s table reports for the same checkpoint. Blue cells mark the pairs of numbers each stage’s change is read from; § 5.2 explains how to read the table.
Figure 4
Figure 3 : SFT training dynamics of the shipped SFT run. (a) Token-level training loss per step (light) and its 50-step moving average (dark). (b) Pass@1 of the checkpoint saved every 100 steps on AIME 25 and AIME 26 ( n=8 , Math-Verify [ 38 ] scoring), GPQA Diamond, and IFEval and IFBench (prompt-level strict). The dashed line marks checkpoint 3799, the one the RL pipeline starts from. The loss steps down at each epoch boundary as the data repeats, but the held-out scores flatten within the first epoch, so later checkpoints buy lower loss without better evaluation.
IFEval
IFBench
GPQA
AIME 25
AIME 26
GLM-4.5-Air
83.00
33.60
73.90
84.20
86.50
Rufus-Air SFT (3799)
88.33
57.75
68.18
90.83
90.00
Δ
+5.33
+24.15
−5.72
+6.63
+3.50
Table 4: SFT-only checkpoint vs. the public GLM-4.5-Air release. All values are pass@1. We compare SFT checkpoint 3799, the checkpoint the RL pipeline starts from (Table 6 ), before any RL, against the public release (RL-included); Δ is the SFT checkpoint minus the public release. The SFT checkpoint already leads on single-turn instruction following and on both AIME years; it trails only on GPQA, one of the gaps the RL stages then target.
Prompts
Prompt tokens
Task
Data source
Verifier
#
%
#
%
Math
HF crawl; synthesized
Math-Verify (canonical match)
57,736
47.7
6.34M
25.3
Science
HF multi-science crawl
Fuzzy string matching
43,199
35.7
4.81M
19.2
Puzzles
Enigmata; ReasoningGym
Generated Python checkers
20,226
16.7
13.91M
55.5
Total
121,161
100
25.07M
100
Table 5: RLVR prompt set. Sources, verifiers, and composition per task family. Math leads on prompt count, Puzzles on token volume—puzzle statements are roughly six times longer. Families are assigned by a keyword heuristic because the crawled problems carry no subject labels. We apply difficulty and correctness filtering (see text) uniformly across families.
Checkpoint
GPQA
AIME 25
AIME 26
Rufus-Air SFT
68.18
90.83
90.00
+ Reasoning RL
73.50
88.02
87.40
Δ
+5.32
−2.81
−2.60
Table 6: Reasoning RL results , measured against the SFT checkpoint the stage trains from (SFT step 3799). Δ is the change contributed by this stage. All values are pass@1 percentages under the same evaluation protocol. GPQA is the stage’s primary reference metric; AIME 25 and AIME 26 are guardrails.
Source
Modality
Problems
% of set
EvolveCoder (synthetic)
functional
14,743
50.1%
Nemotron (competitive)
stdin/stdout
9,393
31.9%
Dolci
mixed †
3,109
10.6%
ADR (algorithmic)
stdin/stdout
2,160
7.3%
Total
53.2% fn / 46.8% i/o
29,405
100%
Table 7: Coding RL data composition after the drop-all-pass screen. All problems are Python. † Dolci splits into 911 functional and 2,198 stdin/stdout problems.
Figure 4 : Coding RL training dynamics. (a) Raw training reward. (b) LiveCodeBench v6 pass@1. The maximum rollout response length is 64K tokens for the first 23 rollout steps and 128K thereafter. Following the extension, the truncation rate falls from 11–28% to below 0.1% . During continued training, reward reaches 0.42 and LiveCodeBench v6 pass@1 peaks at 75.9 at step 33, the checkpoint we carry forward.
Checkpoint
Pass@1
Pass@8
Reasoning RL
68.6
85.1
+ Coding RL
75.9
87.4
Δ
+7.3
+2.3
Table 8: Coding RL results on LiveCodeBench v6, measured against the Reasoning RL checkpoint the stage trains from. Δ is the change contributed by this stage. All values are percentages from one single-stage evaluation at n=8 ; this protocol differs from the one behind Table 1 , so the LiveCodeBench values here are not directly comparable with the headline table.
Figure 5 : IF RL training dynamics . In (a) and (c) the light curve is the per-step value and the dark curve a 5-step moving average. Reward rises throughout without flattening, so the run ends on its step budget rather than on convergence. Response length falls from ∼5.4 K tokens to ∼3.6 K and stays there. The IFBench eval in panel (b) is scored by the in-loop training-time harness, so its values differ slightly from the reported protocol of Table 9 .
Checkpoint
IFEval
IFBench
Multi-challenge
AdvancedIF
GPQA
AIME 25
Coding RL
90.5
63.8
31.1
45.5
67.9
87.6
+ IF RL
94.5
77.8
55.8
62.9
71.1
86.4
Δ
+4.0
+14.0
+24.7
+17.4
+3.2
−1.2
Table 9: Instruction-following RL results , measured against the Coding RL checkpoint the stage trains from. Δ is the change contributed by this stage. IFEval, IFBench and AdvancedIF report prompt-level accuracy.
Figure 6 : General Agent training dynamics , shown up to step 68, the checkpoint the pipeline carries forward. (a) Training reward; (b) held-out AWM score, scored by the in-loop training-time harness; (c) trajectory length in tokens. Light curves are per-step values and dark curves a 5-step moving average.
Checkpoint
MCP-Atlas
Tau2-Retail
AIME 26
IFEval
GPQA
LiveCodeBench
IF RL
34.95
74.00
87.08
94.73
73.36
71.21
+ General Agent
42.75
83.80
87.60
96.16
76.39
73.36
Δ
+7.80
+9.80
+0.52
+1.43
+3.03
+2.15
Table 10: General Agent results , measured against the IF RL checkpoint the stage trains from. Δ is the change contributed by this stage. Tau2-Retail here is pass@1 under the Claude Sonnet 4.5 [ 5 ] user simulator ( n=4 ), so it is comparable within this table but not to the Sonnet 5 [ 6 ] row of Table 1 .
Checkpoint
Terminal-Bench 2.1
SWE-bench Verified
AIME 26
IFEval
LiveCodeBench
General Agent
38.76
65.60
87.60
96.16
73.36
+ Coding Agent
40.17
67.80
87.40
95.68
75.64
Δ
+1.41
+2.20
−0.20
−0.48
+2.28
Table 11: Coding Agent results , measured against the General Agent checkpoint the stage trains from. Δ is the change contributed by this stage. The stage was trained on the compute available rather than to convergence, so these are the values a short run produced.
Figure 7 : Search Agent training dynamics . (a) Training reward, the judge-scored mean over each rollout batch; (b) pass rate on the held-out 200-prompt validation set, evaluated every five steps (the step-4 point is taken from a reference run, because the logged run crashed before its first evaluation finished); (c) tool iterations per episode. Light curves are per-step values and dark curves a 5-step moving average.
Dataset
Original #
Valid #
Avg. Chosen Score
Avg. Rejected Score
Avg. Chosen Tokens
Avg. Rejected Tokens
Arena Human Preference
84,402
53,502
17.06
12.08
1055.1
623.7
HelpSteer3
38,459
29,506
9.60
2.95
439.1
373.2
HH-RLHF
112,052
75,815
-4.23
-8.41
78.8
72.9
Table 13: RLHF prompt datasets. We score each dataset’s chosen and rejected responses with the Skywork reward model; Valid counts the pairs left after we drop missing, empty, and otherwise degenerate responses. Bradley–Terry [ 16 ] scores carry no shared zero across datasets, so read the score columns within a row rather than as grounds for ranking the three.
Figure 8 : RLHF training dynamics of the shipped run, trained on HH-RLHF prompts against the Skywork-Reward-V2-Qwen3-8B reward model with a linear length penalty. (a) Training reward; (b) held-out reward-model score on Arena-Hard v2 prompts, evaluated every five steps; (c) mean response length. Light curves are per-step values and dark curves a 5-step moving average; the dashed line marks iteration 29, the checkpoint shipped as Rufus-Air .
Arena-Hard v2 (HP)
Arena-Hard v2 (CW)
Search Agent
83.06
38.56
+ RLHF
89.05
52.97
Δ
+5.99
+14.41
Table 14: RLHF results. Vanilla RLHF optimizes the raw reward-model score under a linear length penalty (§ 3.8 ); the resulting checkpoint (iteration 29 of the run in Figure 8 ) is the shipped model, Rufus-Air . The reference row is the Search Agent checkpoint the stage starts from, so Δ is the change over the RLHF stage. Both columns are Arena-Hard v2 win rates.
Figure 9 : Agentic rollout stack overview.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Version
Megatron-LM
3714d81
SGLang
24c9100
Strands Agents
1.50.2
PyTorch [ 74 ]
2.9.1 (cu129)
TransformerEngine [ 67 ]
2.10.0
FlashAttention [ 21 ]
2.7.4.post1
Appendix
Table 15: Pinned training and rollout stack. Short hashes are upstream commits.
Stage
Optimizer
LR
Batch
Samples/prompt
Opt. steps/rollout
Response budget
Nodes
SFT
AdamW
5e−5→5e−6
4096 seq
n/a
n/a
128K ctx
64
Reasoning RL
AdamW
1e−6
256 prompts
16
2
30K tok
8
Coding RL
AdamW
1e−6
128 → 64 prompts
64
8
64K → 128K
32
IF RL
AdamW
1.5e−6
256 prompts
16
1
16K tok
8
General Agent
AdamW
1e−6
48 prompts
64
1
32K tok
16
Coding Agent
AdamW
1e−6
32 prompts
32
1
60K tok
16
Appendix
Table 16 : Per-stage training configuration summary, covering all eight stages of the pipeline.
Dataset
Terms
Released-data provenance
Ring-lite-sft-data [ 54 ]
Apache-2.0
Aggregates open datasets including BigMath, DeepScaleR, and DAPO, whose problems originate from public competition and textbook sources.
hermes_reasoning_ tool_use [ 40 ]
Apache-2.0
Derived from ShareGPT conversations originally collected from sharegpt.com .
ToolMind [ 111 ]
Apache-2.0
Draws on xLAM, When2Call, glaive-function-calling-v2, ToolACE, and BUTTONInstruct; released responses were generated with DeepSeek-V2-Chat, Mixtral-8x22B-Instruct, and DeepSeek-V3.
tool-use-multiturn-reasoning [ 39 ]
Apache-2.0
DeepSeek-R1 and QwQ-32B generated the released data.
ToolMind-Web-QA
Apache-2.0
QA pairs follow rules derived from Wikipedia entity–relation graphs; search trajectories were generated by MiroThinker using Qwen3-235B-A22B-Thinking-2507.
Toucan-1.5M [ 108 ]
Apache-2.0
Qwen3-32B, Kimi-K2, and GPT-OSS generated the released data.
Appendix
Table 17 : Public datasets used in the SFT mixture. We transcribe licenses and usage restrictions from the corresponding dataset pages; upstream generators and source corpora describe the released datasets rather than Rufus-Air-side response regeneration.
Qwen3.5
GPT-OSS
Ring-flash-2.0
Solar-Open
Sarvam
Mistral-Small-4
122B-A10B
120B
100B-A6B
102B-A12B
105B-A10B
119B-A7B
IFBench
[ 81 ]
[ 81 ]
–
[ 101 ]
–
[ 63 ]
IFEval
[ 81 ]
[ 81 ]
–
[ 73 ]
[ 86 ]
[ 123 ]
Multi-challenge
[ 81 ]
[ 87 ]
–
[ 101 ]
–
–
Arena-Hard v2 (HP)
[ 69 ]
[ 69 ]
–
–
–
–
AIME 25
[ 69 ]
[ 71 ]
[ 53 ]
[ 73 ]
[ 86 ]
[ 63 ]
Appendix
Table 18 : Source of every publicly reported cell in Table 1 . Each entry is the reference the value was taken from, coloured by the kind of source: orange , the model’s own card or technical report; blue , a baseline reported in another model’s technical report; navy , the benchmark owner’s or a third-party leaderboard. A dash marks a cell that Table 1 leaves empty.
Dataset
Checkpoint
Iters
Search
Scrape
Python
Iters corr./wrong
BrowseComp
GLM-4.5-Air
54.0
46.6
7.3
0.0
32.9 / 61.6
Coding Agent
70.6
47.6
22.1
0.9
42.8 / 85.1
+ Search Agent
75.3
51.9
22.2
1.0
42.5 / 92.0
Seal-0
GLM-4.5-Air
18.0
11.3
6.4
0.3
17.9 / 18.0
Coding Agent
31.6
15.3
15.5
0.8
26.9 / 36.3
+ Search Agent
29.8
13.8
14.9
1.0
25.0 / 34.4
Appendix
Table 19: Mean tool usage per trajectory under the evaluation harness, for the public GLM-4.5-Air release and the Search Agent stage’s starting and final checkpoints (Table 12 ). Iters counts assistant turns with at least one tool call; the last column splits it by the judge’s verdict. The GLM-4.5-Air BrowseComp row covers 950 of the 1,266 tasks; the other rows pool 2–5 evaluation rounds per benchmark.
Domain
User simulator
Rufus -Air †
GLM-4.5-Air
INTELLECT-3
Nemotron-3
Sonnet 4
74.0
69.5
65.0
78.5
Sonnet 4.5
76.0
74.0
70.5
75.5
Airline
Sonnet 5
80.5
77.0
76.0
86.0
Sonnet 4
69.1
57.0
68.4
68.2
Sonnet 4.5
84.6
74.8
73.7
82.2
Retail
Sonnet 5
86.6
80.7
75.7
86.4
Appendix
Table 20 : Tau2-Bench pass@1 under three user simulators; Sonnet 5 is the setting used in Table 1 . Only the customer simulator model changes down a block. † An earlier checkpoint than the final Rufus-Air in Table 1 .
The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL directly to intermediate pre-training checkpoints. We find that RL is effective very early, and often matches the full SFT→RL pipeline early as well. Through experiments on harder problems, we find that targeted pre-training data composition is a strong lever for RL effectiveness, even more so than model scale. Beyond reasoning accuracy, applying RL directly to base checkpoints expands the model's distribution; the sharpening effect reported in recent work arises only when RL follows SFT. The general capabilities of the model remain essentially unchanged by RL, while they degrade following SFT. Finally, we merge RL and SFT objectives by parallel averaging, which outperforms across all other training methods discussed, across metrics, while preserving general capabilities. Together, these results suggest that LLM training might benefit from an expanded use of RL.
Post-training has become a crucial step for unlocking the capabilities of large language models, with reinforcement learning (RL) emerging as a critical paradigm. Recent RL-based post-training has increasingly split into two paradigms: reinforcement learning from human feedback (RLHF), which optimizes models using human preference signals in target domains, and reinforcement learning from verifiable rewards (RLVR), which operates in verifier-backed environments. The latter has dominated recent reasoning-oriented post-training because it delivers stronger gains and higher efficiency on domain-specific tasks (e.g., reasoning). However, although in-domain RL training achieves promising performance, it still requires a substantial amount of GPU compute, which remains a major barrier to broad adoption. In this work, we study the generalization ability of RLHF learned from scratch from a small set of interactions in open-ended environments, and investigate whether the conversational abilities it explicitly acquires can implicitly transfer to downstream tasks such as mathematical reasoning and code generation, namely GRLO. Specifically, on Qwen3-4B-Base backbone, GRLO improves the average performance across all domains from 24.1 to 63.1 with only 5K prompts and 22.7 GPU hours, requiring about 46× less data and 68× less compute than a strong in-domain RLVR baseline. The resulting model is even competitive with Qwen's released post-trained models which required a much larger training cost. Notably, a subsequent in-domain RLVR stage brings only selective gains, mainly on harder competition-math benchmarks. We hope GRLO offers a simple and efficient recipe for building broadly capable post-trained models. Our code and data will be available at: https://github.com/SJY8460/GRLO.
Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization. This work presents NebulaExp, a fully transparent, ablation-driven post-training pipeline built on Qwen3-8B-base, covering two orthogonal model branches: general instruct model and complex reasoning-specialized model. We curate a raw corpus of 3.84M multi-source SFT samples and a 200K verifiable RL candidate pool, and design an end-to-end data processing stack including response distillation, multi-dimensional cross-verification filtering, fine-grained difficulty grading, task classification and diversity-aware sampling. For the Instruct branch, our three-stage optimized supervised fine-tuning approach NebulaExp-Ins-SFT improves the average benchmark score from the 55.01 baseline of Qwen3-8B-nothink to 60.99. GRPO reinforcement learning then further elevates the average score to 61.85. For the Reasoning branch, medium-difficulty GRPO RL improves average reasoning score from 73.88 to 75.17. To address RL's dependency on task verifiers, we systematically investigate single-teacher and multi-teacher OPD (MOPD): utilizing merely 4K instruction-following samples and outperforms RL baseline by 3.26 points on IFEval with +4.43 average overall gain; MOPD fuses four domain-specialist teachers with merely 10K samples, lifting average performance by 4.18 over the base model. This report provides a fully reproducible empirical post-training recipe for 8B-scale LLMs, and comprehensively dissects the capability trade-offs among instruction adherence, mathematical reasoning, code generation and general knowledge.