Research
The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
Overview Research area: Efficient reasoning in large language models — specifically parallel test-time compute (Best-of-N sampling) and the "overthinking" phenomenon in reasoning models. Technical lev

- arXiv
- 2510.21067
- Published
- 2025-10-24
- Authors
- Raul Cavalcante Dinardi, Bruno Yamamoto, Anna Helena Reali Costa, Artur Jordao
AI summary
Overview
Research area: Efficient reasoning in large language models — specifically parallel test-time compute (Best-of-N sampling) and the "overthinking" phenomenon in reasoning models.
Technical level: Intermediate. The core idea is simple enough for newcomers, but the paper leans on concepts such as chain-of-thought, Reinforcement Learning with verifiable rewards, reward normalization in GRPO/GSPO/PPO, and sentence-embedding distance analysis.
Scope: The paper proposes and validates a single heuristic — pick the shortest of N sampled solutions — as a cheaper alternative to self-consistency for reasoning models on math and coding benchmarks.
What This Paper Is About
Reasoning models spend many tokens "thinking," and prior work has shown they often keep generating after they should have stopped, wasting compute. This paper argues the waste also happens when the model is unsure: longer chains-of-thought are systematically more hedged and more likely to be wrong, so length itself carries a usable signal about correctness. The authors therefore test the counterintuitive rule of selecting the shortest solution from N parallel samples, on the theory that this avoids the model's verbose "overthinking" behavior at lower cost than scoring-based methods.
Key Contributions
- Demonstrates that selecting the shortest solution out of N samples is a competitive and cheaper alternative to self-consistency for parallel test-time compute, and frames this as a Pareto improvement in the accuracy-versus-token-usage trade-off.
- Provides evidence that reasoning models operate in two regimes — a concise "conventional regime" and a verbose "overthinking regime" — and locates a critical point where the overthinking regime begins to dominate, defined as the mode of the overall token-usage distribution.
- Uses linguistic uncertainty markers and embedding-based divergence of 500-word chain-of-thought chunks to show trend breaks at that critical point, supporting the two-regime hypothesis across three models and two benchmarks.
- Offers a theoretical explanation grounded in existing RL training algorithms (GRPO, GSPO, and most PPO implementations), whose reward functions normalize rewards by solution length, giving a model an incentive to dilute negative reward with frivolous extra reasoning.
Main Findings
- Shortest beats longest, and matches self-consistency on AIME. With N=5, DeepSeek-R1 reached 89.0% on AIME with the shortest-solution heuristic versus 89.2% for self-consistency, 85.0% for individual attempts, and 78.2% for the longest solution.
- The same pattern holds for Grok-3-mini and Qwen3-32B on AIME. Grok-3-mini: 85.2% shortest, 86.2% self-consistency, 81.0% individual, 74.9% longest. Qwen3-32B: 92.5% shortest, 93.0% self-consistency, 89.5% individual, 85.5% longest.
- On LiveCodeBench, self-consistency is not applicable because answers are not directly comparable, yet the shortest-solution heuristic still works: DeepSeek-R1 79.2% versus 76.5% individual and 76.5% longest; Grok-3-mini 69.2% versus 69.5% individual and 66.8% longest; Qwen3-32B 79.5% versus 78.6% individual and 76.8% longest.
- Longer responses carry more uncertainty. Across all three models and both benchmarks, when two solutions to the same problem were both correct, the longer one showed a higher density of uncertainty markers per 100 words in 67.0% (DeepSeek-R1, AIME), 67.5% (DeepSeek-R1, LiveCodeBench), 67.4% (Grok-3-mini, AIME), 63.7% (Grok-3-mini, LiveCodeBench), 58.2% (Qwen3-32B, AIME), and 65.8% (Qwen3-32B, LiveCodeBench) of cases.
- A trend break appears at the critical point. Solution length correlates positively with uncertainty markers below the critical point; above it, the relationship breaks down, which the authors read as a genuine regime change.
- Reasoning paths diverge up to the critical point, then plateau. Average pairwise cosine distance between embeddings of same-position 500-word chain-of-thought chunks rises until the critical point and stays roughly constant afterwards.
- Shortest-solution selection works with N=2, whereas self-consistency needs at least three solutions for a consensus to exist, making the heuristic usable in tighter compute budgets.
- Efficiency comes from early stopping. Once one of the parallel solutions finishes, the others are terminated; under synchronous generation, the discarded candidates would have been longer, hence worse by the paper's hypothesis.
- Accuracy times token usage grows sublinearly, which the authors show through Pareto curves comparing accuracy against token usage.
Methodology in Plain English
The authors took three reasoning models — DeepSeek-R1, Grok-3-mini, and Qwen3-32B — and ran each at temperature 1 on a random subset of 400 AIME competition questions and on LiveCodeBench v5, sampling five solutions per problem with minimal, answer-extractable prompts. For every problem they compared four selection rules: individual attempts, shortest solution, self-consistency (majority answer), and longest solution.
To explain why the shortest answer works, they did two kinds of text analysis on the chain-of-thought. First, they counted hedging words from a fixed list (for example "maybe," "wait," "perhaps," "however," "hold on," "let me reconsider," "i made an error") and measured their density per 100 words, then compared that density between shorter and longer correct solutions to the same problem. Second, they split each chain-of-thought into 500-word chunks, embedded every chunk with the all-MiniLM-L6-v2 sentence transformer, and measured the cosine distance between chunks occupying the same position in two different solutions to the same problem. Both analyses were plotted against solution length relative to a "critical point," which they defined as the mode of the overall token-usage distribution — a proxy for where conventional-regime solutions peak and overthinking starts to take over.
Why This Matters
Impact on research. The paper reframes overthinking as something that happens not only on easy problems where a model should have stopped, but also on hard problems where the model is uncertain. It also links a training-side cause (length-normalized rewards in GRPO, GSPO, and most PPO implementations) to an inference-side selection rule, and it challenges the assumption that Best-of-N requires a complex scorer or a dedicated reward model. Because the heuristic only needs comparable lengths rather than comparable answers, it extends test-time scaling to tasks where self-consistency cannot be applied at all — the paper explicitly states this is the case for LiveCodeBench in this study.
Real-world applications:
- Code generation assistants, where outputs are programs rather than a single extractable answer and majority voting is ill-defined.
- Mathematical and quantitative tutoring or solving tools that must trade accuracy against per-query cost.
- Any deployed reasoning system with a strict token or latency budget, since early stopping on the first completed solution cuts the compute of the remaining candidates.
- Batch or agentic pipelines that already fan out multiple candidates per prompt and need a cheap, model-agnostic way to choose among them.
Industry relevance. Serving reasoning models is token-dominated, so a selection rule that improves the accuracy-per-token curve without training a reward model or adding a scoring pass is directly relevant to inference cost. The paper states the method gives better theoretical end-to-end latency and can discriminate with as few as two samples, which matters for cost-constrained deployments. The authors do not report measured latency or dollar figures, so any savings claim beyond token accounting is not quantified in the paper.
Future Directions
- Test whether the two-regime structure and the shortest-solution heuristic hold outside mathematics and coding, especially on tasks the paper itself uses as the motivating case for non-comparable outputs.
- Determine whether changing the length-normalization term in GRPO, GSPO, or PPO reward functions reduces or removes the overthinking regime, and whether the shortest-solution heuristic becomes less necessary as a result.
- Investigate whether the critical point can be predicted per model and per problem set rather than estimated empirically from the mode of the observed token-usage distribution.
- Compare the heuristic against alternatives such as dedicated reward models and other Best-of-N scorers on the same Pareto axis, since the paper benchmarks it only against self-consistency, individual attempts, and longest-solution selection.
- Establish how sensitive the results are to sampling settings; the paper reports temperature = 1 and N = 5 (with discussion of N = 2), but does not report results across other temperatures or seeds.
Target Audience
Researchers and engineers working on inference-time scaling, efficient reasoning, or deployment of reasoning models will benefit most, followed by practitioners who need a low-cost selection rule for generated code or numeric answers. Readers interested in the training dynamics of RL-with-verifiable-rewards models will find the reward-normalization argument relevant, though the paper presents it as an explanation grounded in cited prior work rather than as a new derivation.
Authors’ abstract
Reasoning models represent a significant advance in LLM capabilities, particularly for complex reasoning tasks such as mathematics and coding. Previous studies confirm that parallel test-time compute-sampling multiple solutions and selecting the best one-can further enhance the predictive performance of LLMs. However, strategies in this area often require complex scoring, thus increasing computational cost and complexity. In this work, we demonstrate that the simple and counterintuitive heuristic of selecting the shortest solution is highly effective. We posit that the observed effectiveness stems from models operating in two distinct regimes: a concise, confident conventional regime and a verbose overthinking regime characterized by uncertainty, and we show evidence of a critical point where the overthinking regime begins to be significant. By selecting the shortest answer, the heuristic preferentially samples from the conventional regime. We confirm that this approach is competitive with more complex methods such as self-consistency across two challenging benchmarks while significantly reducing computational overhead. The shortest-answer heuristic provides a Pareto improvement over self-consistency and applies even to tasks where output equality is not well defined.