Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.
Figures & tables
Figure 1: Training dynamics of Qwen3.8-Flash-Next under FP4 rollout. Conventional FP4 QAT exhibits increasing train-rollout discrepancy and degraded test score on Terminal-Bench 2.1, while TRACE maintains stable alignment and enables the policy to progressively adapt to FP4 rollout.
Figure 2: Overview of TRACE for FP4 RL (FP4 W/A + FP4 KV) Training of MoE Models .
Figure 3: Two Examples of Train–Rollout Discrepancy under Different FP4 RL Methods.
Figure 4: Problem Analysis of FP4 RL Methods on Qwen3.5-35B-A3B.
Figure 5: Data Communication Pipeline of Rollout-Side Quantization Information.
Figure 6: Difference in FP4 Codebook Entries Between Training and Rollout Activations.
Qwen3.5-35B-A3B
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
NVFP4 W/A + NVFP4 KV
QAT
55.1
70.7
63.4
49.2
59.6
QaRL
54.1
73.3
57.9
49.2
58.6
QUADS
60.4
80.5
75.4
59.0
68.8
TRACE
66.4 ( ↑ 6.0 )
86.3 ( ↑ 5.8 )
78.5 ( ↑ 3.1 )
70.0 ( ↑ 11.0 )
75.3 ( ↑ 6.5 )
Table 1: Performance ( ↑ ) comparison of QAT, QaRL, QUADS, and TRACE under joint NVFP4 W/A/KV rollout setting on Qwen3.5-35B-A3B for reasoning RL tasks. The best performance is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Model
BF16
QAT
QUADS
TRACE
Qwen3.5-122B-A10B
33.4
28.8
29.1
33.0 ( ↑ 3.9 )
Qwen3.8-Flash-Next
68.8
60.4
66.4
70.6 ( ↑ 4.2 )
Qwen3.8-2.4T-A95B
90.3
88.7
85.9
90.2 ( ↑ 1.5 )
Table 2: Performance ( ↑ ) on larger-scale MoE LLMs, including Qwen3.5-122B-A10B and Qwen3.8-Flash-Next (125B-A6B) trained on coding RL tasks and evaluated on DeepSWE and Terminal-Bench, respectively. Qwen3.8-2.4T-A95B is trained on long-horizon RL task and evaluated on GDPval. The best score in each row is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Figure 7: Adaptation for NVFP4 (W/A+KV) with TRACE on HMMT25 during RL training.
Figure 8: Training dynamics of Qwen3.5-35B-A3B on reasoning RL tasks, including reward, response length, and test scores on HMMT, AIME24, AIME25, and LiveCodeBench.
Figure 9: Efficiency analysis of TRACE on Qwen3.5-35B-A3B. Decoding throughput is measured using only the rollout engine (SGLang) on 4 GB200 GPUs. End-to-end RL training step time is measured on the reasoning RL tasks using 72 GB200 GPUs.
Figure 10: Rollout time breakdown for TRACE .
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
NVFP4 W/A + BF16 KV
QAT †
63.5
81.7
78.8
60.4
71.1
QUADS †
64.8
83.3
80.4
62.9
72.9
TRACE
67.1 ( ↑ 2.3 )
83.1
82.1 ( ↑ 1.7 )
69.4 ( ↑ 6.5 )
75.4 ( ↑ 2.5 )
BF16 W/A + NVFP4 KV
QAT
65.3
82.3
79.2
65.8
73.2
TRACE
68.1 ( ↑ 2.8 )
83.0 ( ↑ 0.7 )
80.7 ( ↑ 1.5 )
67.3 ( ↑ 1.5 )
74.8 ( ↑ 1.6 )
Table 3: Ablation of isolated NVFP4 weight and NVFP4 KV on Qwen3.5-35B-A3B under reasoning RL tasks. † marks results quoted from QUADS ( Zhuge et al., 2026 ) . The best performance is marked in bold. The performance gain compared to baselines is marked in green inside bracket.
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
W4A8 + MXFP4 KV
QAT
64.9
75.4
76.9
58.8
69.0
TRACE
67.1 ( ↑ 2.2 )
83.3 ( ↑ 7.9 )
81.2 ( ↑ 4.3 )
68.7 ( ↑ 9.9 )
75.1 ( ↑ 6.1 )
W4A4 + MXFP4 KV
QAT
60.2
74.8
73.2
60.7
67.2
TRACE
65.2 ( ↑ 5.0 )
82.0 ( ↑ 7.2 )
79.8 ( ↑ 6.6 )
67.1 ( ↑ 6.4 )
73.5 ( ↑ 6.3 )
Table 4: Performance ( ↑ ) comparison of stabilizing algorithms with different MXFP4 rollout configurations on Qwen3.5-35B-A3B under the reasoning RL task. The best performance is marked in bold. The performance gain compared to the QAT baseline is marked in green inside brackets.
Rollout Config
Algo.
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16 W/A + BF16 KV
Vanilla
67.1
83.8
81.3
67.5
74.9
FP4 W/A + FP4 KV
QUADS
60.4
80.5
75.4
59.0
68.8
TRACE (R-4bit-L40)
67.0 ( ↑ 6.6 )
86.1 ( ↑ 5.6 )
80.1 ( ↑ 4.7 )
70.5 ( ↑ 11.5 )
75.9 ( ↑ 7.1 )
TRACE (R-3bit-L40)
66.8 ( ↑ 6.4 )
85.9 ( ↑ 5.4 )
81.2 ( ↑ 5.8 )
66.8 ( ↑ 7.8 )
75.2 ( ↑ 6.4 )
TRACE (R-2bit-L40)
67.1 ( ↑ 6.7 )
86.1 ( ↑ 5.6 )
80.0 ( ↑ 4.6 )
67.2 ( ↑ 8.2 )
75.1 ( ↑ 6.3 )
TRACE (R-1bit-L40)
66.4 ( ↑ 6.0 )
86.3 ( ↑ 5.8 )
80.5 ( ↑ 5.1 )
70.0 ( ↑ 11.0 )
75.8 ( ↑ 7.0 )
Table 5: Modular sensitivity study of TRACE with FP4 MoE and FP4 KV on Qwen3.5-35B-A3B under reasoning RL tasks. TRACE (R-1bit-L20) is the default config. The best performance is marked in bold. The performance gain compared to the QUADS baseline is marked in green inside brackets.
Figure 11: Training dynamics of Qwen3.5-35B-A3B on reasoning RL tasks, including the minimum/maximum log-probability difference, reward, entropy loss, response length, and test scores on HMMT, AIME24, AIME25, and LiveCodeBench.
Method
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16
67.1
83.8
81.3
67.5
74.9
TIS+SC
66.7
81.7
81.3
66.7
74.1
TRACE
66.4
86.3 ( ↑ 4.6 )
78.5
70.0 ( ↑ 3.3 )
75.3 ( ↑ 1.2 )
Table 6: Performance ( ↑ ) of SC and TRACE on Qwen3.5-35B-A3B under joint NVFP4 W/A/KV rollout. SC is selected by the highest average across the four benchmarks through 400 steps. The best score in each column is marked in bold.
Method
LiveCodeBench
AIME24
AIME25
HMMT25
Average
BF16
67.1
83.8
81.3
67.5
74.9
Vanilla NVFP4
63.9
75.7
78.4
63.4
70.4
4over6
64.1
83.8
76.7
59.2
71.0
H-Scale
65.5
82.0
77.8
60.1
71.4
TRACE
66.4 ( ↑ 0.9 )
86.3 ( ↑ 2.5 )
78.5 ( ↑ 0.1 )
70.0 ( ↑ 6.6 )
75.3 ( ↑ 3.9 )
Table 7: Final FP4 policy performance ( ↑ ) under different RL training and quantization regimes on Qwen3.5-35B-A3B. For PTQ baselines, RL training is completed with BF16 rollout before applying post-hoc FP4 quantization. In contrast, TRACE uses joint FP4 weight/activation and FP4 KV-cache rollout throughout RL training.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Performance evolution of Qwen3.8-Flash-Next on coding RL tasks: maximum and minimum log-probability differences, mean reward, training entropy loss, mean response length, and test score on Terminal-Bench 2.1.
Figure 13: Training dynamics for Qwen3.8-2.4T-A95B (Max): (a) maximum log-probability difference, (b) mean reward, (c) training entropy loss, and (d) mean response length.
Figure 14: Qwen3.8-Flash-Next output throughput under four expert-weight/KV precision combinations. All settings use one GB200, BF16 PLE offload, 8K-token inputs, and 64 submitted requests.
Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scaling for accuracy preservation with native W4A4 FP4 GEMMs for higher throughput than FP8. However, we find that directly applying NVFP4 to MoE RL rollout is impractical. NVFP4 rollout with BF16 training collapses after roughly 150 steps, accompanied by rapidly growing rollout-trainer log-probability gaps. Through training-inference error analysis and controlled ablations, we identify activation error, rather than weight error, as the dominant source of FP4 RL instability: weights can be synchronized and aligned by a shared quantization-dequantization path, whereas activations are recomputed online and error is amplified by the coarse E2M1 grid. Therefore, to stabilize NVFP4 RL for MoE, we propose QUantization-error Alignment across Dual Sides (QUADS). On the trainer side, we introduce Asymmetric Quantization-Aware Training fake-quantizing weights while keeping activations unquantized for better alignment. On the rollout side, Residual Activation Compensation corrects high-error activation channels while preserving native W4A4 GEMMs. In our MoE RL experiments on several benchmarks, QUADS achieves BF16-level accuracy, improves average pass@1 by 21.49 points over naive NVFP4 RL, and delivers ~16% higher rollout throughput than FP8.
We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout-training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul -- a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout's compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format -- whose three-level hierarchical scaling preserves resolution under FP4's tight 4-bit budget -- closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to-end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.
Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints. Microscaled FP4 formats enable efficient FP4 deployment; however, fully quantizing weights, activations, and KV caches (W4A4KV4) causes severe reasoning degradation that existing PTQ and QAT fail to recover. We identify that FP4 failures concentrate on low-entropy tokens--precise symbolic commitments such as digits and operators--where quantization noise inflates sampling errors that cascade through reasoning traces. Based on this insight, we propose ReQAT, a reasoning-centric FP4 training framework with three components: (i) Trace-Aligned QAT (TAQ), which revisits identical reasoning traces to focus updates on critical low-entropy decisions; (ii) Selective Entropy Minimization (SEM), which reinforces confidence at low-entropy positions; and (iii) Q-FIT, a quantization-friendly initialization that jointly calibrates RoPE-consistent KV cache transformations to stabilize QAT. Under the same training budget, ReQAT not only recovers but surpasses BF16 fine-tuning accuracy, while delivering up to 3.9x throughput speedup on NVIDIA DGX Spark and 3.1x on B200.
Janghwan Lee, Sihwa Lee, Jinseok Kim +4
1Hanyang University, 2Rebellions Inc., Republic of Korea.