Research
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
Overview Research area: Reinforcement learning with verifiable rewards (RLVR) for large language model reasoning, specifically token-level credit assignment and exploration. Technical level: Advanced.

- arXiv
- 2609.33781
- Published
- 2026-09-27
- Authors
- Woongyeong Yeo, Minki Kang, Chanuk Lee, Sangwoo Park, Jinheon Baek, Sung Ju Hwang
AI summary
Overview
Research area: Reinforcement learning with verifiable rewards (RLVR) for large language model reasoning, specifically token-level credit assignment and exploration.
Technical level: Advanced. The paper combines an empirical resampling study, an entropy-based advantage reweighting rule, and two theoretical propositions grounded in KL-regularized optimization.
Scope (one sentence): The paper proposes Entropic Advantage Policy Optimization (EAPO), a method that reweights GRPO's response-level advantage across tokens using normalized policy entropy coupled with the advantage sign, so that high-entropy decisions are reinforced in successful responses and low-entropy decisions are penalized in failed responses.
What This Paper Is About
RLVR trains LLMs using a single reward per response, so every token in a trajectory receives the same learning signal even though individual reasoning decisions contribute differently to the outcome. Prior attempts at finer-grained credit assignment need auxiliary models, extra sampling, or privileged information; policy entropy is attractive because it is free, but existing entropy-based methods treat uncertainty identically under success and failure. This paper asks whether uncertainty should guide reinforcement and penalization the same way, and answers no by designing an asymmetric reweighting scheme.
Key Contributions
-
An empirical characterization of entropy–outcome asymmetry. Using continuation resampling on Qwen3-4B/8B-Base, the authors show that success reached at high-entropy positions is less repeatable ("surprising success") while failure at low-entropy positions recurs ("repeated failure"), and that high-entropy positions in incorrect responses still retain paths to recovery.
-
EAPO, an entropy-guided credit assignment method. EAPO normalizes token entropies with batch-level 10th and 90th percentiles, then couples the signed uncertainty
sign(Â^i)·h_{i,t}with an exponential reweighting that redistributes the response advantage across completion tokens while preserving its within-response mean. It requires no auxiliary models, token-level supervision, or extra rollouts. -
A theoretical interpretation. Proposition 1 recasts the reweighting as the unique solution to a KL-regularized problem over token positions, where uniform credit is the reference allocation and
κcontrols how strongly that anchor is relaxed. Proposition 2 analyzes negative-advantage updates, showing that stronger penalties monotonically shrink the sampled-token probability and increase KL distortion among unchosen alternatives — motivating penalty attenuation at uncertain positions. -
Broad empirical validation. EAPO is evaluated on six competition mathematics benchmarks, on eight out-of-distribution Reasoning Gym tasks, and through analyses of training dynamics, pass@k coverage, answer diversity, and epistemic-marker frequency.
Main Findings
-
Success under uncertainty is less repeatable; confident failure recurs. On 20 moderately difficult math problems per model with 64 continuations per window, originally correct Qwen3-4B-Base responses dropped from 0.578 to 0.390 reward when resampled from high- versus low-entropy windows (Δ = −18.8), while originally incorrect responses rose from 0.070 to 0.163 (Δ = +9.3). Qwen3-8B-Base showed the same pattern (0.625 → 0.418, Δ = −20.7; 0.032 → 0.159, Δ = +12.7). Continuation overlap was far lower at high-entropy windows (4.1 and 7.3 for Qwen3-4B-Base, versus 29.1 and 27.8).
-
Token-level language mirrors these dynamics. Across 2,376 Qwen3-4B-Base responses, exploratory expressions such as "let's" and "consider" were enriched at high-entropy positions in correct responses, conclusion markers such as "finally" and "confirm" at low-entropy positions in incorrect responses, and exploratory expressions such as "try" and "instead" at high-entropy positions in incorrect responses.
-
Best overall accuracy on six math benchmarks. EAPO attained the highest mean avg@32 and pass@32 for all four backbones. With Qwen3-4B-Base and Qwen3-8B-Base it reached mean accuracies of 31.0% and 34.0%, exceeding the strongest entropy-based baselines (EntropyAdv and 80/20, respectively) by 5.6 and 4.3 percentage points, and outperforming RLRT on both metrics without privileged information. With the reasoning backbones it reached 72.4% (Qwen3-4B) and 74.3% (Olmo-3-7B-Think-DPO).
-
Generalization beyond mathematics. On eight Reasoning Gym tasks, EAPO achieved the highest macro-averaged scores of 34.89% and 43.02% for Qwen3-4B-Base and Qwen3-8B-Base, improving on the best baselines (80/20 and HAPO) by 1.46 and 3.02 percentage points.
-
Broader coverage under larger sampling budgets. Pass@k on AIME26 with Qwen3-4B-Base, swept from 1 to 256 samples, was highest for EAPO at every budget; at k = 256 EAPO solved an additional problem that no baseline solved.
-
Greater answer diversity on hard problems. On problems with avg@32 ≤ 25% under the initial model with N = 32 samples, EAPO recorded normalized answer entropy of 0.751 and a collision rate of 0.101, versus 0.657/0.166 for GRPO, 0.659/0.159 for EntropyAdv, 0.647/0.173 for HAPO, 0.672/0.161 for 80/20, and 0.659/0.168 for RLRT.
-
Highest rate of exploratory cues. EAPO produced the highest frequency of epistemic markers per 1,000 generated tokens across all six mathematical benchmarks, by a substantial margin over GRPO and the other exploration-oriented methods.
-
The asymmetry is what matters. In a nine-cell ablation over reinforcement preference
b+and penalization preferenceb-in {−1, 0, +1} on Qwen3-4B-Base, fixing high-entropy reinforcement (b+ = +1) and shifting penalization from high to low entropy improved accuracy by 5.64 percentage points. The EAPO configuration(b+, b-) = (+1, −1)scored 31.00 / 61.27 (avg@32 / pass@32), 6.52 points above uniform credit (24.48 / 50.43), while the reversed configuration(−1, −1)scored lowest at 22.56 / 47.83. -
Entropy concentration helps, with a stability trade-off. Ablating
κon AIME25/AIME26, performance rose from 10.42/36.67 and 7.50/33.33 at κ = 0 to 17.08/46.67 and 15.73/50.00 at κ = log 4; κ = log 8 reached 18.23 on AIME25 avg@32 and κ = log 16 reached 16.15 on AIME26. Larger κ improved faster but destabilized training, so κ = log 4 was chosen as default. -
EAPO does not win by inflating entropy or length. Entropy-based RLVR methods increased response length, and EAPO did too, but extending baseline responses to comparable mean lengths did not close the gap. EAPO ended with lower average entropy than EntropyAdv and HAPO while retaining higher entropy than GRPO in later training.
Methodology in Plain English
The work proceeds in three stages.
First, a diagnostic. The authors take math problems on which Qwen3-4B/8B-Base produces both correct and incorrect answers, locate matched 32-token windows of high and low mean token entropy inside those responses, and resample 64 continuations from the prefix before each window. If high-entropy decisions are genuine branching points, resampling from them should diverge more; if low-entropy decisions are hardened habits, resampling from them should reproduce the same outcome. Both predictions hold, and the direction differs by whether the original response was right or wrong. A complementary word-frequency analysis over the top and bottom 10% entropy tails of 2,376 responses shows which kinds of words sit at each extreme.
Second, a reweighting rule. Standard GRPO gives every token the same response-level advantage. EAPO instead computes each token's entropy under the old policy, rescales it to [0, 1] using the 10th and 90th percentiles of entropies within the rollout batch, and multiplies that normalized value by the sign of the response advantage. An exponential function of this signed uncertainty produces per-token weights, normalized by their within-response mean so the total credit is redistributed rather than inflated. The consequence is structural: for positive advantages, high-entropy tokens get more weight; for negative advantages, low-entropy tokens get more weight; and the weights stay positive so the sign of the advantage is never flipped. The entropies are detached from the computation graph, so gradients flow only through the policy objective.
Third, validation. EAPO is dropped into the existing GRPO-style objective on DAPO-Math-17k-Processed with a DAPO-style configuration, trained on Qwen3-4B-Base, Qwen3-8B-Base, Qwen3-4B, and Olmo-3-7B-Think-DPO, and compared against GRPO, EntropyAdv, HAPO, 80/20, and RLRT using avg@32 and pass@32. The authors then probe coverage across sampling budgets, answer diversity on low-accuracy problems, epistemic-marker frequency, and the two design knobs (κ and the sign–entropy pairing).
Why This Matters
Impact on research. Most entropy-based RLVR methods apply the same uncertainty preference regardless of whether the response succeeded or failed. This paper shows that this shared preference is measurably suboptimal: the ablation moves from 22.56 to 31.00 avg@32 purely by flipping the allocation direction between reinforcement and penalization. It also supplies a theoretical bridge — a KL-anchoring argument for why the reweighting is principled, and a single-step analysis of why penalizing uncertain tokens distorts the distribution over alternatives. The result matters methodologically because EAPO achieves this without auxiliary reward models, extra sampling, or privileged information, and the diagnostic resampling protocol is reusable for studying exploration in other settings.
Real-world applications (potential, not evaluated in the paper):
- Training reasoning assistants that must produce reliable multi-step mathematical or logical derivations.
- Improving coverage in code generation and algorithmic problem solving, where the Reasoning Gym tasks (Graph, Cube, Sudoku, Family, K&K, Anagrams, Zebra, Palindrome) serve as proxies.
- Test-time scaling pipelines, where higher pass@k at large sampling budgets translates into fewer wasted samples on problems a model can actually solve.
- Fine-tuning pipelines constrained by compute or data-privacy budgets, since the method needs no auxiliary models or privileged labels.
Industry relevance. EAPO is a drop-in change to the advantage term inside an existing policy optimization objective, with a single new hyperparameter κ. It requires no new infrastructure, no reward-model training, and no extra rollout generation, which makes it cheap to adopt relative to process-reward or teacher-conditioned approaches. The reported gains hold across base models and already-strong reasoning models, which matters for teams that fine-tune post-trained checkpoints rather than pretrained bases.
Future Directions
-
Dependence on the model's own uncertainty calibration. The authors note that entropy reflects learned preferences and confidence, so the credit allocation inherits the model's biases; the method's reliability depends on how well that uncertainty tracks each decision's actual contribution to the outcome. Characterizing where this alignment breaks down is left open.
-
Extending the mechanism analysis across states. The current analysis examines alternative continuations from a single reasoning state; how findings combine across different states or full trajectories is unexplored.
-
Iterative discovery settings. The limitations section points toward maintaining diverse candidate solutions and iteratively reusing or combining their useful components, citing recent work in scientific discovery, as a direction for strengthening exploration.
-
Choosing
κmore systematically. Largerκimproves faster but destabilizes training, and the default κ = log 4 was picked by observed stability rather than by a stated rule, leaving a principled schedule or selection criterion as an open question.
Target Audience
This paper is for researchers and engineers working on post-training and reinforcement learning for LLM reasoning, particularly those interested in credit assignment, exploration, and entropy-based policy shaping. It is also relevant to practitioners seeking low-overhead improvements to GRPO-style pipelines, and to readers studying uncertainty signals in language models. Comprehension of GRPO, policy entropy, and KL-regularized objectives is assumed; the empirical sections on resampling and diversity are accessible without that background.
Authors’ abstract
Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.