Research
SeLaR: Selective Latent Reasoning in Large Language Models
SeLaR: Selective Latent Reasoning in Large Language Models Overview Research area: Natural Language Processing — reasoning in large language models, specifically training-free latent reasoning as an a

- arXiv
- 2604.08299
- Published
- 2026-04-09
- Authors
- Renyu Fu, Guibo Luo
AI summary
SeLaR: Selective Latent Reasoning in Large Language ModelsOverview
Research area: Natural Language Processing — reasoning in large language models, specifically training-free latent reasoning as an alternative to Chain-of-Thought (CoT) decoding.
Technical level: Intermediate. The paper assumes familiarity with token embeddings, decoding distributions (top-k, top-p, temperature), and entropy, but the core ideas are described in accessible terms.
Scope: The paper proposes SeLaR, a training-free framework that activates soft-embedding ("latent") reasoning only at high-uncertainty decoding steps and applies a contrastive regularization to keep multiple reasoning trajectories alive, evaluated on five reasoning benchmarks across three Qwen3 model scales plus one DeepSeek-R1-Distill model.
What This Paper Is About
Chain-of-Thought reasoning forces a model to commit to one discrete token at every step, discarding information about other plausible continuations. Latent reasoning methods replace those discrete tokens with soft embeddings — probability-weighted mixtures of token embeddings — but existing training-free versions apply this globally, perturbing steps where the model is already confident, and the soft embeddings tend to collapse back toward the single highest-probability token. SeLaR's goal is to decide when latent reasoning should be activated and how to keep it from collapsing into single-path behavior.
Key Contributions
-
An empirical observation about uncertainty structure: The authors show that during CoT decoding, step-wise normalized entropy follows a long-tail pattern, with most steps clustered in a low-entropy (deterministic) region and a sparse tail of high-entropy (exploratory) steps. They demonstrate that activating latent reasoning only at those exploratory steps substantially outperforms global activation.
-
An entropy-gated selective mechanism: SeLaR computes normalized entropy over the top-k tokens at each step and uses a threshold τ to choose between standard discrete decoding (low entropy) and soft-embedding latent reasoning (high entropy), leaving confident steps untouched.
-
An entropy-aware contrastive regularization: To counter the documented "premature collapse" of soft embeddings toward the dominant token, SeLaR pushes the soft embedding away from the top-1 token's direction by an amount proportional to the normalized entropy.
-
A cost-effectiveness evaluation protocol: The paper introduces Tokens per Correct Answer (TPCA), a metric that combines accuracy with token counts on correct and incorrect samples, and shows SwiR's apparent token savings on correct samples are an artifact of survivorship bias.
Main Findings
- Consistent average gains across model scales: SeLaR achieves the highest average accuracy on all three Qwen3 scales, improving over CoT (Sampling) by +0.75%, +3.88%, and +2.45% on Qwen3-1.7B (60.96% to 61.71%), Qwen3-8B (79.68% to 83.56%), and Qwen3-32B (82.38% to 84.83%) respectively. It is the only method that consistently surpasses CoT at every model size.
- Largest gains on the hardest benchmarks: On Qwen3-8B, SeLaR raises AIME 2024 from 76.67% to 83.33% (+6.66%) and AIME 2025 from 66.67% to 80.00% (+13.33%). The authors attribute this to entropy gating concentrating latent reasoning on the most consequential steps while contrastive regularization prevents collapse exactly there.
- Baselines are inconsistent: Soft Thinking and SwiR occasionally match or beat CoT on individual benchmarks, but their average performance frequently falls below CoT (for example, SwiR averages 76.40% vs. CoT Sampling's 79.68% on Qwen3-8B).
- Better cost-effectiveness than SwiR: On TPCA, SeLaR outperforms SwiR by 6.5, 4.8, 52.4, and 27.2 percentage points on GSM8K, MATH500, AIME 2024, and AIME 2025. On AIME 2024, SeLaR reduces TPCA by 19.2% relative to CoT while SwiR inflates it by 33.2%. The sole exception is GPQA, where SeLaR's TPCA is 9.7% higher than CoT.
- Survivorship bias in token-efficiency claims: SeLaR's token count on correctly-answered samples falls within −6.0% to +1.6% of CoT across all benchmarks, indicating no runtime overhead. SwiR's apparent −19.9% reduction on AIME 2024 is attributed to answering only the easier 60% of problems correctly; TPCA reveals its true +33.2% cost inflation.
- Both components are necessary: Removing selective activation drops Qwen3-8B average from 83.56% to 78.37% (a 5.19% drop, below the CoT baseline). Removing contrastive regularization drops it to 75.74% (a 7.82% drop), with AIME 2024 falling from 83.33% to 70.00% and AIME 2025 from 80.00% to 60.00%.
- Contrastive regularization preserves multiple trajectories: A logit-lens analysis over N=200 branching steps from 10 random AIME 2025 problems on Qwen3-8B (k=10) shows that without regularization, top-1 overlap rises from roughly 0.45 to roughly 0.73 while top-2 overlap stagnates near 0.40. With regularization, top-2 overlap climbs to roughly 0.60 and top-1 settles near 0.55 — both substantial — indicating trajectories coexist rather than one simply replacing the other.
- Latent reasoning is rare: Exploratory steps account for 6.2%–13.8% of total reasoning tokens on Qwen3-8B, averaging 10.0%. AIME 2024 (τ=0.4) shows the highest activation frequency and GPQA-Diamond (τ=0.7) the lowest.
- Robust to hyperparameters: With k=3, performance is stable across τ ∈ [0.3, 0.7], with the best average accuracy (80.86%) at τ=0.5. Smaller k is better: k=3 gives 80.86% versus 76.66% for k=5 and 76.76% for k=7.
- Cross-family results are weaker: On DeepSeek-R1-Distill-Llama-8B, SeLaR reaches the highest average accuracy (60.53%), outperforming CoT (Sampling) by 2.77% and SwiR by 1.25%, but gains are less pronounced than on Qwen3. The authors note this model shows higher activation frequencies (8.8%–15.3%) than Qwen3-8B (6.2%–13.8%) and less confident models trigger excessive exploratory steps.
- Illustrative case study: On an AIME 2025 geometry problem, standard CoT computes Arc HJ = 23° at the critical exploratory step and produces an incorrect answer of 334, while SeLaR computes Arc HJ = 24° and produces the correct answer of 336.
Methodology in Plain English
The authors start from a measurement: they look at how uncertain a model is at each decoding step, measured as entropy over the top-k candidate tokens (renormalized and divided by log k to sit in [0,1]). They find the distribution of these values is long-tailed — most steps are confident, a few are ambiguous.
SeLaR then uses that entropy as a switch. At every step it computes the normalized entropy over the top-k tokens. If it is at or below a threshold τ, the model uses its normal discrete token as the next input, exactly as standard CoT does. If it is above τ, the model instead feeds a soft embedding — a probability-weighted mixture of the top-k token embeddings — as the next input. This keeps confident steps stable and reserves exploration for ambiguous ones.
Because soft embeddings are known to drift toward the single most probable token (collapsing back to greedy-like behavior), SeLaR adds a second step at exploratory steps only. It computes the direction from the dominant token's embedding to the soft embedding, and adds an amount of that direction scaled by the normalized entropy. High uncertainty produces a stronger push away from the dominant token; as confidence returns, the push fades naturally. Nothing is trained or fine-tuned.
Evaluation uses Qwen3-1.7B, Qwen3-8B, and Qwen3-32B on GSM8K, MATH500, AIME 2024, AIME 2025, and GPQA-Diamond, with all methods using temperature 0.6, top-p 0.95, top-k 20, and min-p 0.0, run on 4× NVIDIA RTX PRO 6000 GPUs. Baselines are CoT with sampling, CoT with greedy decoding, Soft Thinking, and SwiReasoning (SwiR). Dataset-specific thresholds are τ=0.6 for GSM8K, τ=0.5 for MATH500, τ=0.7 for GPQA-Diamond, τ=0.4 for AIME 2024, and τ=0.5 for AIME 2025, with k=3 fixed.
Why This Matters
Impact on research: The paper reframes training-free latent reasoning as a question of when to intervene, not just what to substitute. Its entropy long-tail observation gives a principled, cheap signal for selective activation, and its logit-lens protocol offers a way to test whether any latent method genuinely preserves multiple trajectories or merely trades one dominant token for another. The TPCA metric also challenges the common practice of reporting token counts on correct samples alone.
Real-world applications:
- Mathematical and competition problem solving, where the paper's largest gains appear (AIME 2024 and AIME 2025)
- Multi-step symbolic and quantitative workflows such as engineering calculations or financial modeling that resemble MATH500-style tasks
- Any deployment where reasoning cost per query matters, since SeLaR applies latent reasoning to only about 10% of tokens on average
- Knowledge-intensive question answering, where the paper notes the benefit is smallest (GPQA's TPCA rose 9.7% versus CoT)
Industry relevance: SeLaR is training-free and lightweight, meaning it can be added as a decoding-time modification on top of existing models without parameter updates or retraining pipelines. The quantified token cost per correct answer gives practitioners a concrete basis for deciding whether such a method is worth deploying relative to plain CoT.
Future Directions
- Move beyond the input token embedding space. The authors state that operating at the input embedding level is inherently less expressive than manipulating hidden states directly, and call for future latent reasoning work in the hidden-state space.
- Address sensitivity to base model confidence. SeLaR helps confident models (for example Qwen3-8B) more than less confident ones (for example DeepSeek-R1-Distill-Llama-8B), which trigger exploration too often. Confidence-aware activation or signals beyond entropy are proposed as remedies.
- Develop better activation signals. The current trigger is token-level entropy with a dataset-specific threshold; whether a single learned or adaptive criterion could replace per-dataset thresholds is left open.
- Broaden evaluation. The paper notes that gains are dataset-dependent and notably weaker on knowledge-intensive GPQA, suggesting that where latent exploration actually pays off needs more systematic characterization.
Target Audience
Researchers and engineers working on LLM reasoning and decoding-time interventions will benefit most, particularly those interested in latent reasoning, Chain-of-Thought alternatives, and inference efficiency. Practitioners evaluating whether to add a training-free reasoning enhancement to an existing model will find the TPCA analysis and ablation results directly useful. Readers without background in token distributions and decoding strategies will need some grounding, but the paper's central framing — activate exploration only when the model is genuinely uncertain — is intuitive on its own.
Authors’ abstract
Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressiveness of discrete token sampling. Recent latent reasoning approaches attempt to alleviate this limitation by replacing discrete tokens with soft embeddings (probability-weighted mixtures of token embeddings) or hidden states, but they commonly suffer from two issues: (1) global activation injects perturbations into high-confidence steps, impairing reasoning stability; and (2) soft embeddings quickly collapse toward the highest-probability token, limiting exploration of alternative trajectories. To address these challenges, we propose SeLaR (Selective Latent Reasoning), a lightweight and training-free framework. SeLaR introduces an entropy-gated mechanism that activates soft embeddings only at low-confidence steps, while preserving discrete decoding at high-confidence steps. Additionally, we propose an entropy-aware contrastive regularization that pushes soft embeddings away from the dominant (highest-probability) token's direction, encouraging sustained exploration of multiple latent reasoning paths. Experiments on five reasoning benchmarks demonstrate that SeLaR consistently outperforms standard CoT and state-of-the-art training-free methods.