Research
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
Overview Research area: Efficient reinforcement learning (RL) post-training of large language models — specifically low-precision (FP4) quantization for Mixture-of-Experts (MoE) language models. Techn

- arXiv
- 2610.07767
- Published
- 2026-10-06
- Authors
- Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang
AI summary
Overview
Research area: Efficient reinforcement learning (RL) post-training of large language models — specifically low-precision (FP4) quantization for Mixture-of-Experts (MoE) language models.
Technical level: Intermediate to Advanced. The core idea is intuitive, but the paper assumes familiarity with quantization-aware training (QAT), FP4 number formats (NVFP4, MXFP4), KV caches, and RL post-training pipelines such as GRPO.
Scope: The paper introduces TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework that aligns the quantized computation paths used during RL rollout generation and RL training for MoE models, and reports accuracy and efficiency results on four MoE models.
What This Paper Is About
RL post-training of LLMs is expensive because the model must repeatedly generate long trajectories during rollout, motivating low-precision (FP4) rollout. However, FP4's coarse numerical grid creates a discrepancy between the quantized rollout path and the quantized training path, which can cause policy mismatch, degraded RL performance, or training collapse — a problem that is amplified in MoE models where numerical differences can change expert routing. The paper's goal is to reduce that train-rollout discrepancy directly, rather than independently improving the quantization accuracy of each path, so that FP4 rollout preserves BF16-level RL performance while retaining FP4's speed advantage.
Key Contributions
- Rollout-guided quantization-aware training (QAT). Instead of standard round-to-nearest (RTN) on the training side, TRACE uses the FP4 codeword actually produced on the rollout side to choose between the two neighboring normalized FP4 E2M1 codewords bracketing the training-side activation, directly targeting inconsistency in FP4 rounding.
- Efficient quantization-information caching (mantissa-only train-rollout communication). Because rollout guidance creates large data movement, TRACE retains only mantissa and scale information, and only from selected deeper layers of the model, rather than complete rollout-side quantization records.
- Evaluation across four large-scale MoE models and three RL settings. The paper tests Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, Qwen3.8-Flash-Next, and Qwen3.8-2.4T-A95B on reasoning, coding, and long-horizon RL tasks.
- Demonstration that TRACE enables joint FP4 weight/activation plus FP4 KV-cache rollout with performance comparable to BF16 rollout, up to 5.4× rollout speedup, and stronger final FP4 performance than training in BF16 followed by post-hoc FP4 quantization.
Main Findings
-
Joint FP4 rollout matches BF16 on reasoning tasks (Qwen3.5-35B-A3B). Under NVFP4 W/A + NVFP4 KV rollout, TRACE scores an average of 75.3 across LiveCodeBench v6, AIME24, AIME25, and HMMT25, versus 59.6 for QAT, 58.6 for QaRL, and 68.8 for QUADS, and versus 74.9 for BF16 rollout. Per-benchmark TRACE scores are 66.4 (LiveCodeBench), 86.3 (AIME24), 78.5 (AIME25), and 70.0 (HMMT25). The paper states TRACE improves the average score from 68.8 to 75.3, and specifically on HMMT25 from 59.0 to 70.0.
-
Results generalize to larger MoE models and other RL tasks. On Qwen3.5-122B-A10B trained on coding RL and evaluated on DeepSWE v1.1, TRACE reaches 33.0 versus BF16 33.4, QAT 28.8, and QUADS 29.1. On Qwen3.8-Flash-Next (125B total / 6B activated parameters) evaluated on Terminal-Bench 2.1, TRACE reaches 70.6 versus BF16 68.8, QAT 60.4, and QUADS 66.4. On Qwen3.8-2.4T-A95B evaluated on GDPval, TRACE reaches 90.2 versus BF16 90.3, QAT 88.7, and QUADS 85.9.
-
Training stability depends on controlling train-rollout discrepancy. QAT, QaRL, and QUADS show increasingly large train-rollout discrepancy as training proceeds, accompanied by training instability: mean reward decreases substantially at later training stages and test scores degrade sharply. TRACE keeps the train-rollout discrepancy controlled and maintains training dynamics close to BF16 rollout, with reward and downstream test performance remaining stable.
-
The policy progressively adapts to FP4 rollout. Under TRACE, test scores start below the BF16 reference due to quantization-induced degradation, but the gap gradually diminishes; after approximately 120 training steps, TRACE catches up with the BF16 reference and remains comparable, with higher point estimates at several later steps. The paper frames this as adaptation to the low-precision execution environment rather than mere preservation of initial quality.
-
Isolated FP4 weight-only or KV-only quantization is easier than the joint setting. With NVFP4 W/A + BF16 KV, TRACE averages 75.4 (versus QAT 71.1 and QUADS 72.9); with BF16 W/A + NVFP4 KV, TRACE averages 74.8 (versus QAT 73.2). Both isolated settings are closer to the BF16 reference of 74.9 than the joint FP4 configuration, which the paper says highlights the difficulty of joint FP4 rollout.
-
TRACE works across FP4 formats, including MXFP4. Under W4A8 + MXFP4 KV, TRACE raises the average from 69.0 (QAT) to 75.1; under the more aggressive W4A4 + MXFP4 KV, it raises the average from 67.2 (QAT) to 73.5.
-
Only compact rollout-side information is needed. Retaining one mantissa bit across all 40 layers yields an average of 75.8; restricting to the latter 20 layers (the default configuration) yields 75.3; 10 layers yields 74.1; and 5 layers yields 73.4 (all against a QUADS baseline average of 68.8). Largely retaining 2, 3, or 4 bits gives averages of 75.1, 75.2, and 75.9 respectively.
-
Comparison with Score Centering (SC). Under joint NVFP4 W/A/KV rollout, SC (as TIS+SC) has a maximum train-rollout log-probability difference of about 1.8–2.1 and a minimum near −4, far from the BF16 reference, whereas TRACE keeps the difference close to BF16. SC's best evaluated checkpoint (at step 380) averages 74.1 versus 75.3 for TRACE, a gain of 1.2 percentage points.
-
Better than BF16 training followed by post-hoc quantization. PTQ pipelines applied to a BF16-trained policy average 70.4 (vanilla NVFP4), 71.0 (4over6), and 71.4 (H-Scale), compared with 75.3 for TRACE trained with FP4 rollout throughout RL.
-
Efficiency. TRACE achieves up to 5.4× higher decoding throughput than BF16 rollout at 128K output length, while keeping throughput close to vanilla FP4 rollout; the full rollout-guided variant without compression incurs substantial throughput degradation. End-to-end RL step time rises from 664 to 713 (7.4% overhead) versus vanilla joint FP4 rollout. In the rollout-time breakdown, the model forward pass accounts for 86%, weight-side rollout guidance 4%, KV-side rollout guidance 2%, and KV dequantization 8%. Retaining one-bit mantissa plus amortized scale metadata for the latter 20 layers consumes 2048 × 20 × 1.5 = 61,440 bits (7.5 KB) of cached quantization information per generated token, roughly a 6× reduction versus caching full quantized activations.
-
Motivating data-volume analysis. For Qwen3.5-35B-A3B with hidden size 2048, KV head dimension 256, two KV heads, 40 MoE layers, 10 full-attention layers, and a 256K-token maximum response length, assuming 4.5 bits retained per quantized value, one generated token produces approximately 45 KB of activation guidance and 5.6 KB of KV guidance (about 50 KB total). With 4,096 trajectories in one RL step, this is up to approximately 51 TB of rollout-side quantization information per step; at an effective GPU-to-CPU bandwidth of 300 GB/s this takes about three minutes in aggregate, but at a 5 GB/s storage bandwidth it would take roughly three hours serially.
Methodology in Plain English
The researchers start from an observation about how FP4 rounding works. Two nearly identical high-precision values can land on opposite sides of a rounding boundary and become two different FP4 codewords. The paper's worked example: training and rollout activations of 60.24 and 59.76 differ by only 0.48, but after applying the same global and block scales their normalized values become 2.51 and 2.49, which round to 3 and 2 — amplifying the discrepancy from 0.48 to 24. The paper also argues that methods like QUADS, which compensate for rollout-side quantization error with a residual, can reduce per-path quantization error while increasing the train-rollout discrepancy (examples show discrepancies of 0.10 and 0.99 arising in cases where vanilla FP4 would give 0 and 1.0 respectively).
TRACE's fix is procedural. During rollout generation, it records the FP4 quantization outcomes for the routed experts' activations and for the FP4 KV states. During the subsequent QAT phase, for each captured training-side activation it normalizes using the rollout scale, reconstructs a rollout-side reference codeword, identifies the two neighboring FP4 codewords bracketing the training activation, and picks whichever of the two is closer to the rollout codeword instead of applying standard round-to-nearest. The paper shows that under this full-information formulation, the chosen codeword's distance to the rollout codeword is no larger than RTN's under the same candidate set — a local, per-site guarantee that the paper explicitly states does not imply monotonic reduction of full-network discrepancy or policy-level divergence.
To keep this affordable, the paper first measures how much rollout-side information is generated (the ~51 TB per step estimate above) and then compresses it. The compression is motivated by two empirical observations: for over 99% of mismatched quantized values across layers, the training and rollout results differ by only one adjacent FP4 codebook entry; and rounding corrections toward lower FP4 codewords occur predominantly in deeper layers. TRACE therefore communicates only mantissa and scale information from the latter half of the model's layers, cached during rollout and transferred to the training engine to reconstruct a compact reference.
Experimental setup: all methods share the same VeRL, Megatron, and SGLang versions, with VeRL coordinating asynchronous RL training, Megatron as the training backend, and SGLang performing rollout generation, deployed with a disaggregated architecture. Policy optimization uses GRPO, R3 is applied to replay rollout-side expert-routing decisions during training, token-level importance ratios are used with a clipping upper bound of 5, and max_version_diff is set to 3 to bound policy staleness. By default, only routed experts in MoE layers and the KV cache in softmax-attention layers are quantized during rollout; all other modules remain in BF16. Each RL step samples 256×16, 128×16, 64×16, and 64×8 trajectories for Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, Qwen3.8-Flash-Next, and Qwen3.8-2.4T-A95B respectively; maximum response length is 32K tokens for reasoning RL and maximum context length is 256K tokens for coding and long-horizon RL; Qwen3.5-35B-A3B and Qwen3.8-Flash-Next are trained for 400 RL steps, and the other two for 50 steps. Reported metrics are sampled average pass@1 over 8 responses for AIME24, AIME25, and HMMT25, and over 10 responses for LiveCodeBench; Terminal-Bench, DeepSWE, and GDPval scores are averaged over 10, 4, and 3 independent runs per task respectively.
Why This Matters
Impact on research. The paper reframes FP4 RL as a cross-path alignment problem rather than two independent accuracy problems, and provides empirical evidence (the rounding-boundary example and the relative excess discrepancy analysis) that minimizing per-path quantization error does not minimize train-rollout discrepancy. It also reports an adaptation effect — policies progressively recovering quantization-induced degradation during RL — which suggests that low-precision execution can be trained into a policy rather than only bolted on afterwards.
Real-world applications:
- Large-scale RL post-training of MoE reasoning and coding models, where rollout generation is the dominant cost.
- Agentic and long-hor
Authors’ abstract
Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.