Skip to content
AI.info

Research

EasyPPO: Stabilizing the Critic Is Key

Overview Research area: Reinforcement learning for large language model (LLM) post-training, specifically the critic (value network) in Proximal Policy Optimization (PPO). Category on arXiv: Machine L

EasyPPO: Stabilizing the Critic Is Key
arXiv
2609.36802
Published
2026-09-29
Authors
Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez

AI summary

Overview

  • Research area: Reinforcement learning for large language model (LLM) post-training, specifically the critic (value network) in Proximal Policy Optimization (PPO). Category on arXiv: Machine Learning (cs.LG), arXiv:2609.36802v1, published 29 Sep 2026. Authors are affiliated with University of California, Berkeley and Princeton University.
  • Technical level: Intermediate. The paper combines a formal analysis (gradient decompositions, a policy-gradient identity) with three simple, implementable training modifications.
  • Scope: The paper diagnoses two critic failure modes in PPO for LLMs and proposes EasyPPO, a set of three changes evaluated on continuous-reward coding (FrontierCS), binary-reward math reasoning (AIME24), and multi-turn search (Search-R1).

What This Paper Is About

PPO's learned critic is usually presented as a source of stability because it estimates expected returns and reduces policy-gradient variance, but the authors find it is also a major source of instability in LLM reinforcement learning. They identify two specific failure modes: filtering truncated ("overlong") rollouts out of both the actor and the critic shifts the policy objective toward reward conditioned on completion, so truncation can increase even while conditional reward improves; and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates within finite batches. EasyPPO is the proposed fix: actor-only overlong filtering, noise-normalized critic regression, and moderately smaller critic mini-batches.

Key Contributions

  1. Diagnosis of overlong filtering as a policy-objective shift. The authors show that when truncated rollouts are excluded from both actor and critic updates, the critic target changes from E[R | s] to E[R | s, not truncated], and derive the resulting policy gradient as the expected actor-critic gradient conditioned on completion. This conditional reward can improve while overall truncation grows.

  2. Diagnosis of heterogeneous return noise in critic updates. They decompose the critic gradient second moment into a prediction-error term and a return-noise (variance) term, showing that as predictions improve, return noise can dominate and give high-variance prompts disproportionate influence in finite batches.

  3. EasyPPO, three modifications to PPO. (a) Actor-only overlong filtering, so the critic trains on returns from both completed and truncated rollouts; (b) noise-normalized critic regression, weighting each prompt's critic loss by the inverse standard deviation of its sampled returns; (c) moderately smaller critic mini-batches, which confine outlier influence to fewer rollouts during gradient clipping. The standard PPO actor update is retained.

  4. Empirical validation across three task types. EasyPPO is compared against vanilla PPO, VAPO, and HL-Gauss PPO, and is reported as the only compared method stable across all three tasks, with best validation relative gains of 14.89%, 2.28%, and 9.47% over PPO.

Main Findings

  • Filtering strategy matters. On FrontierCS, three strategies were compared: no filtering showed repeated reward collapses; filtering both actor and critic improved reward among non-truncated rollouts but drove the truncation ratio toward 1 with overall reward remaining low; actor-only filtering was the most stable and achieved the highest overall score. Actor-only filtering still showed residual instability, including a spike and reward drop near step 160.

  • Joint filtering changes what the policy optimizes. The analysis shows that joint actor-critic filtering optimizes reward conditioned on completion without directly encouraging the policy to avoid truncation, matching the observed pattern of rising reward among completed rollouts alongside more frequent truncation and low overall reward.

  • Noise normalization balances critic gradients. Offline diagnostics on Qwen3.5-9B at steps 30, 60, and 90 (512 responses per checkpoint, 16 prompts with 32 responses each) showed that without normalization, prompts with more variable returns contribute larger critic gradients, a pattern stronger at later checkpoints. Noise normalization largely removed this dependence, though some gradient outliers remained.

  • Critic mini-batch size involves a tradeoff. An idealized analysis gives an O(m) bound on a single outlier's influence on the average clipped gradient (where m is mini-batch size), versus O(1/m) gradient variance per mini-batch. On FrontierCS, first-crossing steps where explained variance first reached zero decreased approximately linearly with log K, while K = 4 and K = 8 achieved higher explained variance later in training. The default is K = 4.

  • Critic behavior during training. After warmup, EasyPPO's critic gradient norm declined and its explained variance trended steadily upward with relatively small fluctuations, whereas baselines showed sharp swings. HL-Gauss approached an explained variance of 1 only after its actor collapsed.

  • Best validation results on a 0–100 scale. EasyPPO improved over PPO by 1.92, 1.46, and 3.74 points on FrontierCS, AIME24, and Search-R1; gains over the second-best method on each task were 0.91, 1.46, and 0.89 points. EasyPPO remained stable throughout 200–300 training updates in the evaluated runs, and every baseline experienced performance collapse in at least one setting.

  • Seed sensitivity. Three runs each of EasyPPO and PPO + actor-only filtering (the strongest baseline on FrontierCS) were compared on FrontierCS. None of the three EasyPPO runs collapsed, while two of the three actor-only-filtering runs did.

  • Noise normalization enables smaller mini-batches. In the FrontierCS ablation at a fixed critic learning rate of 2×10⁻⁶, without normalization, the performance drop was recoverable with one mini-batch but became a sustained collapse when the same rollout batch was split into four smaller mini-batches. With normalization, both configurations preserved gains.

  • PPO + actor-only filtering is not sufficient. It achieved higher FrontierCS scores than vanilla PPO but still fluctuated substantially and collapsed on AIME; EasyPPO stayed stable with the same filtering rule, attributing the difference to the critic modifications.

Methodology in Plain English

The authors start from a standard PPO setup with verifiable rewards: a prompt is sampled, the policy generates a response, and a scalar reward arrives at the end of the rollout. They simplify the exposition by treating a whole response as one action from the prompt state, while the implementation retains token-level GAE, PPO clipping, and KL regularization, with each rollout batch sampled from the latest policy snapshot (strictly on-policy).

They then analyze two things mathematically. First, they take the critic's population target, E[R | s], and ask what happens when truncated rollouts are dropped: the target becomes E[R | s, not truncated], and they derive the gradient identity showing the policy is being pushed toward completion-conditioned reward. Second, they write the critic gradient second moment for a linear value head and show it equals the feature norm times prediction error plus return variance; dividing the loss at each prefix by the return standard deviation makes the remaining terms scale-free across prompts.

In practice, estimating variance at every token prefix is costly, so they use the prompt's return standard deviation across its token states, estimated from the rollout group (including truncated responses), with weights normalized to mean one across the prompt batch. For discrete rewards with range Δ and group size n, they propose a variance floor ε = Δ/(2√n), so groups with identical returns receive about twice the weight of a group with one maximum and n−1 minimum rewards. They apply the same rollout group for weight estimation and critic training, acknowledging this can introduce bias but works well in practice.

For mini-batches, they split the critic batch into B/m mini-batches, clip each mini-batch gradient to at most c, take one optimizer step, then recompute the next mini-batch gradient at the updated parameters. Training setup: Qwen3.5-9B for FrontierCS, Qwen3.5-9B-Base for AIME and Search-R1; a 30-step critic warmup; rollout batches of 512 responses on FrontierCS and AIME and 1024 on Search-R1; group sizes of 32 on FrontierCS and 16 on AIME and Search-R1; a 32,768-token response limit. FrontierCS starts from an SFT checkpoint trained on 347 nonzero-score trajectories (of 600 sampled by DeepSeek-V3.1 across 200 FrontierSmith problems, after removing 253 zero-score responses, covering 179 distinct problems) for a single epoch. Validation averages over five responses per FrontierCS problem and 32 per AIME24 problem; Search-R1 averages greedy accuracies equally across seven datasets. All methods are implemented in verl.

Why This Matters

  • Research impact: The paper reframes critic instability as a primary obstacle in LLM reinforcement learning rather than an implementation detail, and shows that a well-known recipe (overlong filtering) changes the optimization target when applied to the critic. It also connects variance-weighted regression ideas from classical RL and heteroscedastic supervised learning to the group-sampling structure already present in LLM RL, and it argues that standard-deviation weighting belongs on the critic loss rather than on actor advantages.

  • Real-world applications:

    • Coding and algorithm design agents, where FrontierCS-style continuous scores grade solution quality rather than binary correctness.
    • Mathematical reasoning models, where AIME24-style binary verifiable rewards are used for reasoning post-training.
    • Multi-turn search and tool-use agents, where rewards arrive after multi-step interactions with external systems.
    • GPU kernel and systems optimization, where the paper notes reward scales can differ hugely across problems (roughly 1× speedup for some GEMM solutions versus over 100× for specialized operators).
  • Industry relevance: EasyPPO is a small set of changes on top of an existing PPO implementation in verl, so the practical cost of adoption is low relative to redesigning the training algorithm. Because every baseline collapsed in at least one evaluated setting and EasyPPO did not, the recipe targets the reliability concerns that matter when training long-horizon reasoning models at scale.

Future Directions

  • Reducing weight-estimation bias. The current method uses the same rollout group to both estimate variance weights and train the critic, which the authors note can introduce bias; they suggest alternatives such as offline profiling with online updates.
  • Cheaper variance estimation. Estimating return variance separately at each token prefix is described as costly; the prompt-level approximation and the variance floor could be refined, including for the continuous-reward setting.
  • Better mini-batch and optimizer interaction. The mini-batch study varies clipping granularity and optimizer dynamics together, so the authors note the observed tradeoff is consistent with, but not isolated from, their fixed-parameter analysis.
  • Broader evaluation. The seed-sensitivity study was restricted to two methods on FrontierCS due to training cost, and the full analysis of actor-only filtering is deferred to the appendix, leaving room to test generality across more tasks, models, and filtering regimes.

Target Audience

Researchers and engineers working on reinforcement learning for LLM post-training, especially those implementing PPO-based actor-critic pipelines in frameworks such as verl. It is also relevant to readers interested in continuous-verifiable-reward benchmarks and heterogeneous-reward settings, and to practitioners who need stable long-horizon training runs without late-stage collapse. Some familiarity with PPO, advantage estimation, and critic regression is assumed; the mathematical analysis is self-contained enough for readers with a graduate-level machine learning background.

Authors’ abstract

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.

Read the original paper