Research
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Overview Research area: Large language model reasoning — specifically test-time compute scaling, self-refinement, and selective tool/model querying for small reasoning models (sRMs). Technical level:

- arXiv
- 2609.34327
- Published
- 2026-09-28
- Authors
- Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
AI summary
Overview
Research area: Large language model reasoning — specifically test-time compute scaling, self-refinement, and selective tool/model querying for small reasoning models (sRMs).
Technical level: Intermediate. The paper combines a diagnostic empirical study (counterfactual interventions on intermediate reasoning states) with a training pipeline (SFT plus cost-aware reinforcement learning), so some familiarity with reasoning models, pass@k metrics, and RL fine-tuning helps.
Scope: The paper argues that not all reasoning failures benefit from more thinking, distinguishes "execution bottlenecks" from "knowledge bottlenecks," and trains 4B and 8B models to selectively query stronger external models only when their own parametric knowledge is insufficient.
What This Paper Is About
Small reasoning models are cheap to serve but lag behind larger ones, and a common assumption is that letting them think longer will close the gap. The authors show this assumption is only partly right: further self-reflection mostly consolidates answers the model could already reach, while many failures require information the model simply does not have. Their goal is to teach a small model to diagnose which kind of failure it is facing and, only when needed, ask a stronger external model for targeted help at a chosen cost.
Key Contributions
-
A diagnostic account of reasoning failure. Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the authors show that self-refinement mainly consolidates probability mass onto already reachable solutions rather than making new solutions reachable, separating "execution bottlenecks" from "knowledge bottlenecks."
-
Evidence that help-seeking is a learned capability. The paper finds that smaller models not only possess less parametric knowledge but also use external information less effectively than larger models, so simply being given relevant information is not enough.
-
FlyBy, a selective querying framework with a multi-depth query tool. External models are exposed as a three-depth action inside the reasoning trajectory, letting the model jointly decide whether to query, what to ask, and how much external computation to spend.
-
Trained 4B and 8B models that beat larger models at lower serving cost. FlyBy-4B reaches 45.96% pass@8 on 1,158 hard problems across six benchmarks, surpassing Qwen3-14B (41.64%) at 2.7x lower serving cost, and FlyBy-8B reaches 51.81% pass@8.
Main Findings
-
Self-refinement consolidates rather than discovers. Intervening at 13.4K endogenous epistemic verbalization (EV) occurrences from Qwen3-0.6B/1.7B/4B/8B/14B and Gemma4-E2B/E4B/12B shows that EVs remain common even in smaller models, but their causal effect on value increases substantially with model scale. The limiting factor is not expressing uncertainty but converting reflection into progress.
-
Two distinct failure regimes. Execution-like states (estimated value V(s) > 0, i.e., at least one observed successful continuation) are where reflection helps; knowledge-like states (V(s) = 0) are where reflective cues give only modest gains while problem-relevant oracle information yields substantially larger value gains, and random cues provide little benefit.
-
Knowledge bottlenecks are more common in science than math. Among unresolved states, knowledge-like bottlenecks are substantially more prevalent in scientific reasoning, consistent with greater reliance on specific factual and domain knowledge.
-
Small models are limited on both sides. Information utilization generally improves with model scale, so smaller models both enter knowledge-like states more often and exploit external information less well. The apparent drop at the largest Qwen and Gemma models is attributed to changes in the residual problem set; restricted to shared problems the increasing trend is preserved.
-
Two-stage training works better than prompt-only or single upfront queries. SFT alone reaches 36.39% pass@8; adding cost-aware RL raises it to 45.96% with only a marginal increase in serving cost (7.42 m$ versus 7.15 m$ for the SFT model, on a per-eight-rollout cost basis).
-
Coverage expands beyond the base model's reach. On problems unsolved by Qwen3-4B in 16 no-tool rollouts, FlyBy-4B achieves 28.7% pass@8.
-
Improvement is not just multi-sample luck. FlyBy-4B also exceeds Qwen3-8B at pass@1 (16.85% vs. 15.31%), indicating better individual trajectories rather than only more chances at a correct one.
-
Targeted short queries beat long retrieved contexts. Search-R1 returns much longer contexts per call (2,015 characters on average), while FlyBy-4B uses short external responses (105 output tokens on average) across multiple targeted queries.
-
The advantage over an upfront query grows with difficulty. Query Opening (one query to DeepSeek-V4-Flash before reasoning) is considerably stronger than other baselines, but its gap to FlyBy widens as problems get harder, which the authors attribute to the diminishing utility of a broad, upfront request.
-
The policy queries where thinking will not help. On 50 SuperGPQA problems with 8 rollouts each, executing the selected query increases state value by 5.30 pp on average, whereas suppressing the query and applying budget forcing until another query attempt or 512 additional tokens yields essentially no improvement.
-
RL turns one-shot delegation into iterative acquisition. Compared with SFT, the RL model maintains substantial query probability at later turns, and adaptive depth allocation attains performance close to fixed depth-3 querying at a cost close to fixed depth-1 querying.
-
Economic robustness. Sweeping GPU and API prices, FlyBy-4B remains optimal across a broad range of price regimes, so its advantage is not tied to a particular pricing assumption.
-
Scaling the cognitive core helps. FlyBy-8B reaches 51.81% pass@8, 5.85 percentage points above FlyBy-4B and 10.17 points above Qwen3-14B, while remaining cheaper to serve, consistent with larger models benefiting more from given information.
-
The zero-success classification is a finite-sample criterion. In a robustness check on 160 states with 64 new null continuations each, 28 of the 76 states with no success in the first eight samples became positive by N = 64, yet those reclassified states retained a large intervention gap: oracle information still helped substantially while EV cues did not.
Methodology in Plain English
The authors first run a diagnostic study rather than building a system immediately. They treat a model's partial reasoning trace as a "state," then probe that state by appending a short cue and sampling multiple continuations. A state's "value" is the fraction of those continuations that end at the correct answer, and its "answer entropy" measures how varied the answers are, so productive reasoning moves toward high value and low entropy. To locate moments of self-doubt, they use a fixed nine-word lexicon of hedging expressions such as "wait" and "hmm" (epistemic verbalizations). Crucially, every intervention is paired with a counterfactual from the same state — the same prefix with the lexicon banned, or with a null cue — so the measured effect isolates the intervention rather than differences between problems.
They compare three injected cues: an epistemic cue asking the model to reconsider, an oracle information cue generated by DeepSeek-V4-Pro that supplies a problem-specific hint without revealing the answer, and a random cue drawn from an unrelated problem that preserves the form of extra information but removes relevance.
Building on the finding that relevant information helps most where reflection does not, they train FlyBy from Qwen3-4B and Qwen3-8B. The model can, at any point in its trace, either keep reasoning or invoke a multi-depth query tool. Depth 1 uses DeepSeek-V4-Flash with up to 128 output tokens, depth 2 uses DeepSeek-V3.2 with up to 512, and depth 3 uses DeepSeek-V4-Pro with up to 1,536, at increasing API prices. To prevent trivial delegation, the external model never sees the original problem, and queries with excessive n-gram overlap with the problem are rejected during both training and evaluation.
Training has two stages. Supervised fine-tuning on a small curated set of "rescue" trajectories — cases where one query turns a failed continuation into a success — plus targeted supervision for query generation and post-observation integration, mixed with general reasoning data. Then 80 steps of reinforcement learning with a modified GRPO objective in which cost is penalized only for successful trajectories, so the policy is never rewarded for failing cheaply.
Evaluation covers 1,158 hard problems from six benchmarks, where "hard" means vanilla Qwen3-4B achieves pass@1 of at most 0.25. The main metric is pass@8, and serving cost combines local GPU inference cost with external API cost, converted to USD using measured throughput and third-party prices.
Why This Matters
Impact on research: The paper reframes test-time scaling from a question of "how much compute" to "what kind of compute does this state require." Its paired-intervention methodology gives a reusable way to attribute reasoning failures to execution versus knowledge limits, and its finding that help-seeking must be learned — rather than prompted — challenges the assumption that giving small models access to tools or stronger models is sufficient.
Real-world applications:
- Cost-sensitive deployment of reasoning assistants, where a small on-device or self-hosted model handles most problems locally and pays for external help only on genuinely knowledge-bound problems.
- Scientific and medical question answering, where the paper finds knowledge bottlenecks dominate and where ChemBench, MedXpertQA, and GPQA-Diamond results show large gains (FlyBy-4B reaches 70.73% on ChemBench and 35.37% on MedXpertQA).
- Agentic systems with multiple backends, where a cheap model acts as the "cognitive core" that routes hard sub-problems to stronger, more expensive models.
- Retrieval-augmented pipelines, where the authors' comparison suggests that short, targeted, reasoning-informed queries can outperform returning long retrieved contexts.
Industry relevance: Serving cost is reported in milli-dollars per eight rollouts, and the paper explicitly sweeps GPU and API prices. FlyBy-4B's advantage over Qwen3-14B comes at 2.7x lower serving cost, and FlyBy-8B outperforms Qwen3-14B by 10.17 points while remaining cheaper to serve — a direct argument for smaller cores with selective escalation.
Future Directions
- Reducing reliance on external models. The authors state as a limitation that FlyBy depends on external models, raising practical concerns about availability, privacy, and reliability, and point to open directions in the appendix.
- Generalizing the query tool across backends. The paper reports further experiments on backend generalization, tool leakage, and price fluctuations but leaves these as appendix material rather than main claims.
- Refining the execution-versus-knowledge distinction. The zero-success criterion is finite-sample based; the robustness check shows 28 of 76 zero-success states became positive with 64 continuations, so sharper ways of distinguishing practical accessibility from true reachability remain open.
- Closing the small-model information-utilization gap. Since utilization improves with scale, it is an open question how to make smaller models integrate the external information they receive as effectively as larger ones.
Target Audience
Researchers and engineers working on reasoning-model efficiency, test-time compute scaling, and tool- or model-augmented agents. It is especially relevant to practitioners who deploy small open-weight models under cost constraints and want to decide when escalating to a stronger model is worth paying for, and to researchers interested in mechanistic analysis of intermediate reasoning states and the design of cost-aware RL objectives.
Authors’ abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.