Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.
Figures & tables
Figure 1: Training dynamics of Qwen3.8-Flash-Next under FP4 rollout. Conventional FP4 QAT exhibits increasing train-rollout discrepancy and degraded test score on Terminal-Bench 2.1, while TRACE maintains stable alignment and enables the policy to progressively adapt to FP4 rollout.
Figure 2: Overview of TRACE for FP4 RL (FP4 W/A + FP4 KV) Training of MoE Models .
Figure 3: Two Examples of Train–Rollout Discrepancy under Different FP4 RL Methods.
Figure 4: Problem Analysis of FP4 RL Methods on Qwen3.5-35B-A3B.
Figure 5: Data Communication Pipeline of Rollout-Side Quantization Information.
Figure 6: Difference in FP4 Codebook Entries Between Training and Rollout Activations.
Qwen3.5-35B-A3B
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
NVFP4 W/A + NVFP4 KV
QAT
55.1
70.7
63.4
49.2
59.6
QaRL
54.1
73.3
57.9
49.2
58.6
QUADS
60.4
80.5
75.4
59.0
68.8
TRACE
66.4 ( ↑ 6.0 )
86.3 ( ↑ 5.8 )
78.5 ( ↑ 3.1 )
70.0 ( ↑ 11.0 )
75.3 ( ↑ 6.5 )
Table 1: Performance ( ↑ ) comparison of QAT, QaRL, QUADS, and TRACE under joint NVFP4 W/A/KV rollout setting on Qwen3.5-35B-A3B for reasoning RL tasks. The best performance is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Model
BF16
QAT
QUADS
TRACE
Qwen3.5-122B-A10B
33.4
28.8
29.1
33.0 ( ↑ 3.9 )
Qwen3.8-Flash-Next
68.8
60.4
66.4
70.6 ( ↑ 4.2 )
Qwen3.8-2.4T-A95B
90.3
88.7
85.9
90.2 ( ↑ 1.5 )
Table 2: Performance ( ↑ ) on larger-scale MoE LLMs, including Qwen3.5-122B-A10B and Qwen3.8-Flash-Next (125B-A6B) trained on coding RL tasks and evaluated on DeepSWE and Terminal-Bench, respectively. Qwen3.8-2.4T-A95B is trained on long-horizon RL task and evaluated on GDPval. The best score in each row is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Figure 7: Adaptation for NVFP4 (W/A+KV) with TRACE on HMMT25 during RL training.
Figure 8: Training dynamics of Qwen3.5-35B-A3B on reasoning RL tasks, including reward, response length, and test scores on HMMT, AIME24, AIME25, and LiveCodeBench.
Figure 9: Efficiency analysis of TRACE on Qwen3.5-35B-A3B. Decoding throughput is measured using only the rollout engine (SGLang) on 4 GB200 GPUs. End-to-end RL training step time is measured on the reasoning RL tasks using 72 GB200 GPUs.
Figure 10: Rollout time breakdown for TRACE .
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
NVFP4 W/A + BF16 KV
QAT †
63.5
81.7
78.8
60.4
71.1
QUADS †
64.8
83.3
80.4
62.9
72.9
TRACE
67.1 ( ↑ 2.3 )
83.1
82.1 ( ↑ 1.7 )
69.4 ( ↑ 6.5 )
75.4 ( ↑ 2.5 )
BF16 W/A + NVFP4 KV
QAT
65.3
82.3
79.2
65.8
73.2
TRACE
68.1 ( ↑ 2.8 )
83.0 ( ↑ 0.7 )
80.7 ( ↑ 1.5 )
67.3 ( ↑ 1.5 )
74.8 ( ↑ 1.6 )
Table 3: Ablation of isolated NVFP4 weight and NVFP4 KV on Qwen3.5-35B-A3B under reasoning RL tasks. † marks results quoted from QUADS ( Zhuge et al., 2026 ) . The best performance is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
W4A8 + MXFP4 KV
QAT
64.9
75.4
76.9
58.8
69.0
TRACE
67.1 ( ↑ 2.2 )
83.3 ( ↑ 7.9 )
81.2 ( ↑ 4.3 )
68.7 ( ↑ 9.9 )
75.1 ( ↑ 6.1 )
W4A4 + MXFP4 KV
QAT
60.2
74.8
73.2
60.7
67.2
TRACE
65.2 ( ↑ 5.0 )
82.0 ( ↑ 7.2 )
79.8 ( ↑ 6.6 )
67.1 ( ↑ 6.4 )
73.5 ( ↑ 6.3 )
Table 4: Performance ( ↑ ) comparison of stabilizing algorithms with different MXFP4 rollout configurations on Qwen3.5-35B-A3B under the reasoning RL task. The best performance is marked in bold. The performance gain compared to the QAT baseline is marked in green inside brackets.
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
FP4 W/A + FP4 KV
QUADS
60.4
80.5
75.4
59.0
68.8
TRACE (R-4bit-L40)
67.0 ( ↑ 6.6 )
86.1 ( ↑ 5.6 )
80.1 ( ↑ 4.7 )
70.5 ( ↑ 11.5 )
75.9 ( ↑ 7.1 )
TRACE (R-3bit-L40)
66.8 ( ↑ 6.4 )
85.9 ( ↑ 5.4 )
81.2 ( ↑ 5.8 )
66.8 ( ↑ 7.8 )
75.2 ( ↑ 6.4 )
TRACE (R-2bit-L40)
67.1 ( ↑ 6.7 )
86.1 ( ↑ 5.6 )
80.0 ( ↑ 4.6 )
67.2 ( ↑ 8.2 )
75.1 ( ↑ 6.3 )
TRACE (R-1bit-L40)
66.4 ( ↑ 6.0 )
86.3 ( ↑ 5.8 )
80.5 ( ↑ 5.1 )
70.0 ( ↑ 11.0 )
75.8 ( ↑ 7.0 )
Table 5: Modular sensitivity study of TRACE with FP4 MoE and FP4 KV on Qwen3.5-35B-A3B under reasoning RL tasks. TRACE (R-1bit-L20) is the default config. The best performance is marked in bold. The performance gain compared to the QUADS baseline is marked in green inside brackets.
Figure 11: Training dynamics of Qwen3.5-35B-A3B on reasoning RL tasks, including the minimum/maximum log-probability difference, reward, entropy loss, response length, and test scores on HMMT, AIME24, AIME25, and LiveCodeBench.
Method
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16
67.1
83.8
81.3
67.5
74.9
TIS+SC
66.7
81.7
81.3
66.7
74.1
TRACE
66.4
86.3 ( ↑ 4.6 )
78.5
70.0 ( ↑ 3.3 )
75.3 ( ↑ 1.2 )
Table 6: Performance ( ↑ ) of SC and TRACE on Qwen3.5-35B-A3B under joint NVFP4 W/A/KV rollout. SC is selected by the highest average across the four benchmarks through 400 steps. The best score in each column is marked in bold.
Method
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16
67.1
83.8
81.3
67.5
74.9
Vanilla NVFP4
63.9
75.7
78.4
63.4
70.4
4over6
64.1
83.8
76.7
59.2
71.0
H-Scale
65.5
82.0
77.8
60.1
71.4
TRACE
66.4 ( ↑ 0.9 )
86.3 ( ↑ 2.5 )
78.5 ( ↑ 0.1 )
70.0 ( ↑ 6.6 )
75.3 ( ↑ 3.9 )
Table 7: Final FP4 policy performance ( ↑ ) under different RL training and quantization regimes on Qwen3.5-35B-A3B. For PTQ baselines, RL training is completed with BF16 rollout before applying post-hoc FP4 quantization. In contrast, TRACE uses joint FP4 weight/activation and FP4 KV-cache rollout throughout RL training.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Performance evolution of Qwen3.8-Flash-Next on coding RL tasks: maximum and minimum log-probability differences, mean reward, training entropy loss, mean response length, and test score on Terminal-Bench 2.1.
Figure 13: Training dynamics for Qwen3.8-2.4T-A95B (Max): (a) maximum log-probability difference, (b) mean reward, (c) training entropy loss, and (d) mean response length.
Figure 14: Qwen3.8-Flash-Next output throughput under four expert-weight/KV precision combinations. All settings use one GB200, BF16 PLE offload, 8K-token inputs, and 64 submitted requests.