Research
Sharpening Tax in Post-Training
Sharpening Tax in Post-Training Overview Research area: Reinforcement learning (RL) post-training of large language models, with a focus on multi-turn agentic tasks (tool calling, environment interact

- arXiv
- 2610.01509
- Published
- 2026-10-01
- Authors
- Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li
AI summary
Sharpening Tax in Post-TrainingOverview
- Research area: Reinforcement learning (RL) post-training of large language models, with a focus on multi-turn agentic tasks (tool calling, environment interaction) rather than math or coding.
- Technical level: Advanced. The paper assumes familiarity with RL post-training (PPO, GRPO), pass@k metrics, temperature sampling, and Bayesian posterior estimation, and it includes formal theorems.
- Scope: A systematic study of whether post-training sharpens a base model's behavior, introducing a diagnostic metric ("Sharpening Tax") to measure the coverage loss, plus a training-time sampler (PTGS) designed to reduce it.
What This Paper Is About
Prior work suggested that RL post-training mostly sharpens a base model: it boosts single-shot accuracy (pass@1) but reduces solution coverage (pass@K). That evidence came almost entirely from math and coding, where pre-training already provides enormous exposure. This paper asks whether the same tension holds in agentic tasks, where the assumption has been that tool use and long-horizon coherence must be learned during post-training. The authors find it does hold — and they build a single-number metric to quantify how much test-time scalability post-training gives up.
Key Contributions
- A systematic accuracy-vs-coverage study on agentic tasks. The authors evaluate base and post-trained LLM pairs on multi-turn agentic benchmarks and show that harness-equipped base models can match or surpass their post-trained counterparts in coverage (pass@K) given enough rollouts.
- Sharpening Tax, a diagnostic metric. They propose a scalar that summarizes post-training's effect on test-time scalability, computed as the deficit between base and post-trained policies in raw area scalability A(K) and calibrated scalability S(K).
- Empirical characterization of the tax across scale. Across 42 model–benchmark combinations (14 backbones, three agentic benchmarks, four families), the tax is prevalent, can be estimated cheaply from a few rollouts, and correlates with other evaluation metrics.
- Posterior-tempered group sampling (PTGS). A plug-and-play Bayesian sampler that adapts temperature per prompt based on an online difficulty estimate; applied during RL training (PPO and GRPO), it reduces the tax while improving both pass@1 and pass@K.
Main Findings
- Base models are capable agentic reasoners when given a harness. Pre-trained base models equipped with a lightweight, model-agnostic harness (system prompt plus relaxed tool-calling and parsing interfaces) perform agentic reasoning surprisingly well, and their pass@K curves cross over and surpass post-trained models as the rollout budget grows. For example, on WebShop the base model reaches over 85% pass@128 versus 56% for RL with gemma-4-31B.
- The harness helps base models but hurts post-trained ones. The harness ablation over eight base models shows BFCL pass@1 rising from 5.59 to 15.63 and pass@32 from 15.19 to 49.13; WebShop pass@1 from 6.32 to 9.73 and pass@32 from 39.48 to 49.15; ACEBench pass@1 falls from 44.24 to 42.63 and pass@32 from 89.25 to 87.69. Separately, the same harness hurts post-trained models' performance, consistent with evidence that prompt engineering is not uniformly beneficial for advanced models.
- Larger models cross over earlier. On Gemma-4 scales (4B, 12B, 26B, 31B), the crossover budget in WebShop decreases from k* > 128 for 4B to k* ≈ 3 for 31B.
- Post-training bimodalizes per-task success. Categorizing tasks by outcomes over 128 rollouts, post-training sharply reduces the middle "pass given compute" category (from 87.6% to 30.0% on WebShop with gemma-4-31B) and pushes mass to both extremes: "always pass" grows from 0.0% to 26.0% and "always fail" grows from 12.4% to 44.0%.
- Coverage is traded for consistency. Averaged over 12 model–benchmark pairs (largest backbone per family × three benchmarks), the post-trained policy's pass@K and pass^K curves stay relatively close, whereas the base model's consistency collapses toward zero while its coverage eventually surpasses RL.
- The tax is pervasive and scale-dependent. For the largest backbones, the tax is positive at almost every budget and benchmark; for the smallest backbones it is often negative at small budgets (k ≤ 8), most visibly for Ministral-3-3B, but rises as budget grows. By k = 128, Tax_S(128) > 0 in 36 of the 42 combinations.
- The tax is predictable and informative. Tax_S(8), estimated from 8 rollouts on half the tasks, predicts Tax_S(32) on held-out tasks with Spearman ρ = 0.85. It also correlates with the consistency gap Δpass^8 and is the only early predictor tested that predicts all three future targets with a strong statistical sign.
- PTGS reduces the tax during task-specific RL. Fine-tuning Qwen2.5-7B-Instruct on Sokoban and FrozenLake with PPO and GRPO (200 training steps, means over five runs), PTGS improves both pass@1 and pass@128. On Sokoban, PPO goes from pass@1 46.5 / pass@128 55.0 / Tax_S(128) 0.094 to 61.1 / 69.7 / 0.081 under PTGS; GRPO goes from 36.5 / 55.3 / 0.081 to 39.1 / 72.5 / 0.025. On FrozenLake, PPO goes from 63.7 / 74.1 / 0.039 to 65.0 / 80.0 / 0.020; GRPO goes from 63.4 / 77.8 / 0.029 to 67.3 / 81.2 / 0.028.
- PTGS preserves entropy. In FrozenLake PPO training dynamics, fixed-temperature PPO shows monotonically decreasing output-token entropy converging to a few action sequences, while PTGS keeps entropy much higher and rising repeatedly, ending at the same validation accuracy but with a broader space of solvable problems.
- Sharpening varies by family. Qwen3.5 exhibits the softest sharpening and Gemma-4 the steepest.
- Theoretical results. Proposition 1 gives a probabilistic interpretation of raw scalability A(K) as the expected number of failed attempts before the first success. Theorem 2 shows that if post-training sharpens each task with probability λ to either 0 or 1, then Tax_A(K) = λ · A_Base(K) ≥ 0, and the calibrated scalability ratio depends on λ and on the proportion λ₀ sharpened to 0. Theorem 3 (whose full statement is cut off in the provided content) is described as explaining how PTGS improves the rollout groups RL learns from.
Methodology in Plain English
The authors compare open-source checkpoint pairs — a pre-trained base model and its post-trained counterpart from the same family (for example, gemma-4-31B versus gemma-4-31B-it) — across Gemma-4, Ministral-3, Qwen2.5, and Qwen3.5, spanning 3B to 35B effective parameters. They evaluate on three agentic benchmarks that require multi-turn tool calling: BFCL v4 multi-turn base split, WebShop, and ACEBench. Because intermediate actions and state transitions are checked by design in these environments, a lucky final answer cannot inflate scores in the way it can on math and coding benchmarks.
For each task they draw many independent rollouts (up to 128) and compute three metrics using standard unbiased estimators: pass@1 (accuracy), pass@K (coverage, the probability at least one of k rollouts succeeds), and pass^k (consistency, the probability all k rollouts succeed). They then define raw area scalability A(K) as the sum over k = 1 to K−1 of [pass@K − pass@k], calibrated scalability S(K) = A(K) / [(K−1)(1 − pass@1)], and Sharpening Tax as the base-minus-post-trained difference in either quantity.
To reduce the tax, they propose PTGS: during RL training, for each prompt they maintain discounted cumulative success and failure counts with forgetting factor γ, convert them plus a target success rate p̃ into a Beta posterior, draw a success-probability sample via Thompson sampling, and map it to a temperature T_x = τ^h(p̂_x) — heating up prompts estimated as hard and cooling down prompts estimated as easy. PTGS needs no change to the RL update rule and adds no computational overhead. They test it by training Qwen2.5-7B-Instruct with PPO and GRPO on Sokoban and FrozenLake, with and without PTGS, using τ ∈ [1.2, 1.5], γ = 0.95, and p̃ raised from 0.25 to 0.5 over training.
Why This Matters
- Impact on research: The paper extends the sharpening hypothesis beyond math and coding to agentic domains, where new capabilities were widely assumed to be learned during post-training. It provides a cheap, scalar diagnostic (Sharpening Tax) that is anchored to the base model and predicts other evaluation metrics, giving the field a reusable tool for auditing post-training pipelines. It also challenges the assumption that harnesses and prompt engineering help advanced models uniformly.
- Real-world applications:
- Scientific discovery and automated research: when diversity and creativity matter, running a large base model with many rollouts can beat a post-trained model that has collapsed onto a few behaviors, given a scalable verifier.
- Budget-constrained deployment: the finding that the tax is negative at small budgets (k ≤ 8) for small backbones means the right model choice depends jointly on model scale and test-time budget.
- Model routing: the authors demonstrate routing between base and post-trained models per task using the early tax estimate.
- RL training infrastructure: PTGS plugs into existing algorithms such as PPO and GRPO without modifying the update rule or adding overhead, making it a low-cost drop-in.
- Industry relevance: Modern post-training pipelines prioritize immediate sampling efficiency and consistency. This paper quantifies what that costs in coverage and offers a training-time intervention that recovers part of it, which matters for anyone shipping agentic systems where repeated sampling is used.
Future Directions
- Understanding the sharpening mechanism more precisely. The paper's Theorem 2 models the limiting case where sharpened tasks collapse fully to 0 or 1; how partial sharpening behaves in practice is left open in the provided content.
- Extending PTGS to broader settings. The empirical PTGS evaluation covers Qwen2.5-7B-Instruct on Sokoban and FrozenLake with PPO and GRPO. Whether the gains transfer to larger backbones, other RL algorithms, and more diverse agentic environments is not established here.
- Calibrating the tax for decisions. Since the tax is negative at small budgets for small backbones and positive at large budgets for large ones, how practitioners should set a threshold for choosing between base and post-trained models per deployment is a natural next question.
- Applying early tax estimation at scale. Because Tax_S(8) predicts Tax_S(32) with Spearman ρ = 0.85, using it as a routine evaluation signal — and as a router — is presented as promising but only demoed in a single routing experiment.
Target Audience
Researchers and engineers working on RL post-training, agentic LLM systems, and evaluation methodology. It is most useful for those who design post-training pipelines, choose between base and instruction-tuned checkpoints under test-time compute budgets, or study distribution sharpening and diversity in model outputs. Readers should be comfortable with pass@k notation, RL algorithms, and basic probability.
Authors’ abstract
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.