Skip to content
AI.info

Research

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

Overview Research area: Reinforcement learning (RL) for large language models, specifically low-precision (4-bit) training-and-inference systems and the optimization stability of decoupled sampler/lea

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
arXiv
2610.07043
Published
2026-10-05
Authors
Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang, Yang Yu, Mingfa Feng, Congkai Xie, Shuang Yu, Junjie Lai, Hongxia Yang

AI summary

Overview

Research area: Reinforcement learning (RL) for large language models, specifically low-precision (4-bit) training-and-inference systems and the optimization stability of decoupled sampler/learner RL pipelines.

Technical level: Advanced. The paper combines a first-order dynamical analysis of learner–sampler mismatch with a systems-level implementation of native NVFP4 (W4A4) RL on Blackwell GPUs.

Scope: The paper analyzes why native NVFP4 execution destabilizes GRPO policy optimization and proposes TRIAGE, a direction-aware stabilization method that preserves native 4-bit forward execution on both the sampler and the learner.

What This Paper Is About

In LLM reinforcement learning, responses are generated by a high-throughput inference engine (the sampler) but optimized using probabilities recomputed by a separate training framework (the learner). Running both paths in native 4-bit NVFP4 speeds up rollout but enlarges the discrepancy between the two policies, which can destabilize policy optimization. The paper's goal is to show that this discrepancy must be judged together with the direction of the policy-gradient update — not by its magnitude alone — and to build a stabilization method that controls self-amplifying mismatch while keeping native 4-bit execution.

Key Contributions

  1. A directional characterization of learner–sampler mismatch. The authors derive a first-order expansion of "mismatch energy" showing that tokens with δ·g > 0 amplify existing mismatch while tokens with δ·g < 0 contract it, reducing to the sign interaction A_i·δ_{i,t} for vanilla GRPO, and thereby identifying two amplifying regions 𝒜⁻ = {A_i < 0, δ_{i,t} < 0} and 𝒜⁺ = {A_i > 0, δ_{i,t} > 0}.

  2. Evidence of a directional, locally concentrated precursor to instability. In native NVFP4 GRPO runs, the two amplifying regions become imbalanced before the marginal gap distribution becomes broadly severe, and the affected tail tokens concentrate in a small fraction of response segments that response-level averages hide.

  3. TRIAGE, a selective stabilization mechanism. The method combines segment-level risk diagnosis, direction-selective token gating restricted to 𝒜⁻ tokens, within-response update rebalancing, and a bounded pseudo-Huber repair term on positive-advantage responses.

  4. An end-to-end native NVFP4 RL recipe at scale. Experiments totaling more than 34,500 GPU-hours on B300 GPUs validate native W4A4 execution on both sampler and learner forward paths for dense (Qwen3-4B-Base) and MoE (Qwen3-30B-A3B-Base) models, with up to 1.3× end-to-end RL speedup and BF16-level benchmark performance.

Main Findings

  • The theory: magnitude is not the right target. From the mismatch-energy decomposition ΔD = η(B_amp − B_con + R_cross) + O(η²), a token with δ_{i,t}·g_{i,t} > 0 amplifies the existing gap, whereas δ_{i,t}·g_{i,t} < 0 contracts it. The paper states that "a large mismatch may already be self-correcting, while a moderate one grows systematically when the gradient acts along it."

  • Directional asymmetry appears before magnitude collapse. Across drift-to-pre-terminal windows, the fraction of tokens with |δ| < 0.05 remains at least 83.0% on 4B and 91.9% on 30B, yet the advantage-weighted asymmetry ratio ρ_asym rises from 1.02 to 2.32 by pre-terminal on 4B and from 1.34 to 1.85 on 30B.

  • The terminal stage is a negative-tail explosion. The 5% gap quantile q_{0.05}(δ) reaches −42.75 on 4B and −10.94 on 30B within the respective mismatch estimands.

  • Mismatch is localized, not globally spread, before failure. Using non-overlapping 64-token segments, the global frequency of amplifying-tail tokens decreases from 4.8–4.9% to 2.8–3.6% from healthy to pre-terminal across the two models, yet the top 5% of segments contain approximately one quarter of all amplifying-tail tokens (about five times uniform placement), and adjacent tail-heavy persistence rises from roughly 2.6× to 3.8–4.0× the chance rate.

  • Response averages hide the problem. At pre-terminal, 97.2% (4B) and 79.9% (30B) of responses containing tail-heavy segments still satisfy |δ̄_i| < 0.05. By terminal, roughly half of eligible segments are tail-heavy and top-segment concentration declines.

  • Training stability. Naive NVFP4 becomes unstable after approximately 300 steps on Qwen3-4B-Base with a sharp reward drop and rapid growth in the mean absolute train–rollout log-probability gap; on Qwen3-30B-A3B-Base it shows rapid mismatch growth after roughly 700 steps. TIS delays failure but does not prevent it on 4B, and on 30B its reward trajectory remains below BF16 and TRIAGE in later training. TRIAGE remains stable throughout the evaluated 600- and 1,700-step schedules, approximately doubling the observed stable training horizon on 4B relative to naive NVFP4.

  • Benchmark performance is near BF16. Average across AIME24, AIME25, MATH500, AMC23 and GSM8K: 58.49% versus 58.26% on 4B, and 70.96% versus 72.41% on 30B. At the shared step-300 checkpoint on 4B, TRIAGE outperforms naive NVFP4 and NVFP4+TIS by 5.93 and 2.27 percentage points. NVFP4+TIS scores 63.18% over the full 1700-step run versus 65.87% for the pre-collapse naive NVFP4 checkpoint at step 700.

  • Throughput. With TRIAGE, native NVFP4 achieves 6,745 and 8,717 output tokens/s at global batch sizes 256 and 512, corresponding to 2.27× and 2.30× the BF16 throughput. Iteration time drops from 208.2 s to 172.1 s (batch 256) and from 254.5 s to 195.8 s (batch 512), giving 1.21× and 1.30× end-to-end speedups.

  • TRIAGE overhead is small. Isolating learner cost on eight B300 GPUs with five paired repetitions of the same frozen batch gives a median learner-update overhead of 0.84% and a median absolute increase of 0.225 s.

  • Ablations. On Qwen3-4B-Base, branching all variants from the same full-TRIAGE checkpoint at step 300, full TRIAGE achieves the highest reward; removing repair keeps the mean log-probability gap low but leaves the advantage-driven update imbalance uncorrected. TIS alone remains stable when continued from the TRIAGE checkpoint, unlike its failure from the beginning.

Methodology in Plain English

The authors first formalize the setup: a sampler policy π_s generates rollouts and a learner policy π_l is optimized on them, with the per-token log-probability gap δ_t = log π_l − log π_s. They note that existing corrections such as truncated importance sampling, clipping and masking all act on exp(δ), a magnitude-only statistic, since under synchronous weight transfer the whole off-policy error is carried by that factor.

They then ask what happens to the total mismatch energy D = ½ Σ δ² after one gradient step. A first-order Taylor expansion yields a sum over tokens of κ·δ·g plus a cross-token term, so the sign of δ·g determines whether a token's update amplifies or contracts mismatch locally. The authors are explicit that cross-token coupling, optimizer history and higher-order effects can still change an individual token's realized motion, so they treat δ·g as a local directional risk indicator and use the decomposition at the population level.

They then instrument native NVFP4 GRPO runs of Qwen3-4B-Base and Qwen3-30B-A3B-Base, using four 25-update trailing windows at 25%, 50%, 75% and 100% of each run's characterization horizon (labeled healthy, drift, pre-terminal and terminal). For the method itself, they compute per-segment statistics — a weighted mean gap δ̄_S and a tail fraction f^tail_S — and map them to a gate weight w_S in [w_min, 1]. For negative-advantage responses, only tokens with δ_{i,t} < 0 inside a risky segment are gated to w_S; all other tokens keep weight 1. The gate enters the policy loss as a weighted normalized average, which reallocates update mass rather than uniformly shrinking the whole response. Separately, from a fixed update k_rep onward, a pseudo-Huber penalty is applied to the one-sided violation q_R = [τ_rep − δ̄_R]_+ of repair segments on positive-advantage responses, matching the objective while retaining native NVFP4 W4A4 forward execution on both sampler and learner.

Experiments compare four configurations — BF16, native NVFP4, NVFP4+TIS (truncation threshold C = 2.0) and NVFP4+TRIAGE — with GRPO on the DAPO dataset and a max response length of 20,480, evaluating on AIME 2024, AIME 2025, MATH-500, AMC 2023 and GSM8K, plus training reward, policy entropy and the mean absolute train–rollout log-probability gap.

Why This Matters

Impact on research. The paper reframes low-precision RL stabilization from a problem of minimizing the size of learner–sampler mismatch to a problem of controlling how that mismatch evolves under the policy update. Equation 6 and the amplification/contraction split give an analytically motivated criterion that existing magnitude-based corrections (TIS, clipping, masking) cannot express, and the observed directional asymmetry plus segment-localized tail provides an early-warning signal distinct from the terminal negative-tail explosion.

Real-world applications.

  • Training reasoning models on long mathematical and logical chains, where rollout dominates wall-clock time.
  • Long-horizon coding and agentic RL, where trajectories are long and rollout cost is the primary bottleneck.
  • Serving/inference-side generation for RL data collection at 4-bit precision without restructuring the training stack.
  • Deploying RL fine-tuning on datacenter GPU fleets where 4-bit Tensor Core throughput is otherwise left unused during the training loop.

Industry relevance. NVFP4 is a Blackwell-supported format, and the paper reports runs totaling more than 34,500 GPU-hours on B300 GPUs. A 2.3× rollout throughput advantage with 1.21–1.30× end-to-end speedup and a 0.84% median learner overhead is directly relevant to teams paying for large-scale RL post-training. The paper also notes that end-to-end iteration speedup (approximately 1.3×) remains well below the hardware's acceleration potential because a growing share of runtime goes to parameter synchronization between trainer and sampler and quantization-related reductions such as Amax computation.

Future Directions

  1. Scale beyond the studied regime. The authors state that, due to computational constraints, they do not evaluate models at the 100B+ scale or substantially longer training horizons, and that their main runs are terminated after trajectories show clear convergence or stable late-stage behavior.

  2. Asynchronous RL. The directional criterion is not inherently tied to synchronous sampling and may be especially relevant where policy staleness adds another source of learner–sampler mismatch, but this has not been validated experimentally. Jointly characterizing quantization mismatch and asynchronous staleness is left open.

  3. Closing the systems gap. The paper argues that end-to-end NVFP4 RL efficiency is below the acceleration potential of Blackwell and Vera Rubin Tensor Cores, and suggests co-designing TRIAGE with more tightly integrated low-precision kernels, more efficient parameter synchronization and fused quantization paths.

  4. Finer-grained repair rules. The authors suggest a broader design space beyond the current repair rule, adapting to the magnitude, locality and persistence of directional mismatch rather than treating all mismatch-amplifying tokens uniformly.

Target Audience

Researchers and engineers working on RL post-training for LLMs, low-precision and quantized training systems, and inference/training alignment in decoupled RL pipelines. It is also relevant to systems teams deploying 4-bit formats on NVIDIA Blackwell hardware, and to readers interested in optimization-dynamics analyses of policy-gradient methods. The main text is readable without the appendices, but the derivation in Appendix A and the operating rules in Appendix C require comfort with first-order gradient analysis.

Authors’ abstract

Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. Their tail tokens become concentrated in a small fraction of response segments before mismatch spreads globally. Motivated by these findings, we introduce TRIAGE, a direction-aware stabilization method that uses segment-level diagnosis to selectively rebalance policy-gradient updates and applies bounded repair to residual severe mismatch. TRIAGE modifies the optimization objective while retaining native NVFP4 weight-and activation 4-bit (W4A4) forward execution on both the sampler and learner. Experiments on Qwen3-4B and Qwen3-30B-A3B show stable optimization throughout the evaluated training horizon and achieve full precision level performance across five mathematical reasoning benchmarks, while native NVFP4 with TRIAGE provides up to 2.3x higher rollout throughput than BF16.

Read the original paper