Research
Can Computation from Earlier Problems Help LLMs Solve New Ones?
Overview Research area: Multi-turn large language model reasoning, mechanistic analysis of attention and hidden-state reuse, and parameter-efficient adaptation. Technical level: Advanced. The paper co

- arXiv
- 2609.39394
- Published
- 2026-09-30
- Authors
- Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
AI summary
Overview
Research area: Multi-turn large language model reasoning, mechanistic analysis of attention and hidden-state reuse, and parameter-efficient adaptation.
Technical level: Advanced. The paper combines layerwise representation analysis, nearest-neighbor geometry of hidden states, and a Householder-reflection-based attention modification, though its central question is stated in plain terms.
Scope: The paper asks whether the internal computation left behind by an earlier, independent problem in the same conversation can be captured and re-read to help a frozen LLM solve a later problem, and introduces a 12,288-parameter controller called STAIR to test this.
What This Paper Is About
When a user asks an LLM one problem and then asks another in the same conversation, the earlier problem and its answer remain visible as text. For the model, generating that earlier answer also produced a sequence of internal computations. The paper investigates whether those historical internal states remain useful after the task switches to a new, independent problem, and whether teaching the model how to read them can improve later-turn accuracy. Preliminary experiments show retained history can either raise or lower accuracy depending on the model and task, motivating a method that learns the readout rather than freezing it.
Key Contributions
-
The authors identify a recurring, problem-dependent component in how the current problem's internal representation changes when the source history changes. Isolating this "interaction response" across two disjoint data groups on Qwen3-4B Instruct, they find that the nearest-neighbor structure among current problems is preserved across different histories (48.06% and 46.50% cross-history neighborhood overlap, against 11.73% and 12.08% under a shuffled-identity control), even though this component accounts for only 0.69% and 0.66% of total displacement energy at layer 19.
-
They introduce STAIR (Stale-Token Attention for Inter-query Reuse), which stores attention keys and values captured during earlier response generation in a fixed, read-only bank and learns to redirect current queries when they read that bank during prompt processing (prefill).
-
Across three Qwen models and four benchmarks, STAIR improves mean T2-T4 Avg@4 over the unmodified Native model with history by up to 11.67 percentage points while training only 12,288 parameters in an otherwise frozen backbone.
Main Findings
-
History can help or hurt, depending on model and task. On MATH-500 with Qwen3-4B Instruct, Avg@4 falls from 95.00% under Vanilla (no history) to 93.75% at Native T4. On GPQA-Diamond with Qwen3.5-4B, Avg@4 rises from 63.64% under Vanilla to 75.25% at each of Native T2, T3, and T4.
-
Attention to earlier assistant responses persists deep into the network. At layer 3, earlier assistant-response bodies receive 39.49% and 38.29% of the attention mass in the two data groups; at layer 19 they still receive 12.41% and 12.87%. State displacement also persists into later layers, with a median norm of approximately 17.76 at layer 19 and 62.07 (group 1) and 61.97 (group 2) at layer 35.
-
History changes answers in both directions. Across 500 MATH-500 problems with four samples each (2,000 responses per turn), T4 shows 35 answers changing from wrong to correct and 60 from correct to wrong, for a net Avg@4 change of -1.25 percentage points. T2 shows 37 wrong-to-correct and 47 correct-to-wrong (net -0.50), and T3 shows 40 and 59 (net -0.95).
-
STAIR's largest gains appear on AIME 2025. STAIR raises mean T2-T4 Avg@4 over Native by 11.67 points for Qwen3.5-4B (40.56% to 52.22%), 6.11 points for Qwen3.5-9B (43.33% to 49.44%), and 3.61 points for Qwen3-4B Instruct (34.44% to 38.06%). All three improve Avg@4 at each later turn and mean Pass@4 on this benchmark.
-
Results across benchmarks. On MATH-500, STAIR reaches 94.63% (Qwen3-4B Instruct), 96.27% (Qwen3.5-4B), and 96.78% (Qwen3.5-9B) mean T2-T4 Avg@4; both Qwen3.5 models exceed their Vanilla references. On AMC23† (34 problems after excluding training overlap), Qwen3.5-4B gains 6.37 Avg@4 points over Native, from 68.14% to 74.51%. On GPQA-Diamond, Qwen3-4B Instruct improves both metrics over Native (58.00% to 58.75% Avg@4; 76.94% to 77.78% Pass@4), while the Qwen3.5 models show no comparable gain, and Qwen3.5-9B drops from 79.88% to 75.93% Avg@4.
-
Pooled accuracy by prompt type. Pooling T2-T4 across all three models on MATH-500, AIME 2025, and AMC23†, responses that followed a correct T1 answer (18,879 responses) score 93.96% under Vanilla, 93.42% under Native, and 94.18% under STAIR; responses following a wrong T1 answer (1,425 responses) score 80.07%, 60.35%, and 68.56% respectively.
-
Shared-history control results. Holding T1 fixed on AIME 2025 with Qwen3.5-4B at T2 (30 problems, four matched trajectories each), Native reaches 21.67% Avg@4 and 56.67% Pass@4, while STAIR reaches 48.33% and 80.00%, a 26.67-point Avg@4 gain. A separately trained Bank-free controller reaches 36.67% (+15.00), Direct reflected-read 41.67% (+20.00), Random-reflector 37.22% (+15.56), and K-V pairing permutation 33.06% (+11.39).
-
Weight adaptation and historical access compose. On AIME 2025 with Qwen3-4B Instruct, STAIR uses 12,288 trainable parameters versus LoRA's 4,128,768 (336 times as many) and improves Avg@4 over Native by 3.61 points; adding STAIR to LoRA (4,141,056 parameters total) improves Avg@4 over LoRA by 3.89 points. Pass@4 rises by 4.44 points in both comparisons.
-
Bank length and runtime costs. Restricting the bank to the first 2K, 8K, or 32K saved tokens yields 35.00%, 43.33%, and 39.17% T2 Avg@4, compared with 48.33% for the full bank. Full-bank reading adds a median 0.81 s and 6.60 GiB in peak memory over Native.
Methodology in Plain English
The authors first ran a diagnostic study. Using Qwen3-4B Instruct on MATH-500, they replayed the same current problem after many different three-turn histories while keeping the model weights frozen. Two balanced designs crossed 64 current problems with 128 histories per group, and 128 current problems with 32 histories per group; the three-turn problem-coverage condition alone comprises 8,192 problem-history cells. By measuring attention mass on earlier assistant replies and the shift in the current problem's hidden representation, they showed that history leaves a persistent trace.
To separate what is common to all problems under a history from what is specific to a problem-history pairing, they decomposed each displacement into an overall mean, a problem term, a history term, and a leftover "interaction response." They then checked whether a problem's nearest neighbors in interaction-response space stayed the same when the history changed. Because the additive terms alone reproduce neighborhood overlap at 100%, the interesting result is that the small leftover component retains 48.06% and 46.50% overlap across histories.
STAIR builds on this. During earlier turns, the model's attention keys and values from the response body are captured and stored in an ordered, read-only bank. Keys are captured after key projection and normalization but before rotary position encoding. At the next turn, the current query is passed through a learned Householder reflection, which reverses the query's component along a learned direction without changing its length. The original and reflected queries both read the same frozen bank, and STAIR uses the difference between the two readouts, so the differential weights sum to zero and only the redistribution of attention over stored content is injected. Reflection directions are learned at layers 3, 11, and 19, and the projected update is added to the native self-attention output during prefill only; decoding then follows the original path. Only the reflection directions are trained, on responses the frozen backbone generates for the DAPO-Math-17k dataset (14,806 training and 128 validation problems, one epoch), using a negative log-likelihood loss plus a KL term (lambda = 1) toward the Native model's distribution.
Why This Matters
The paper reframes the residue of completed tasks in a conversation as a resource whose readout can be learned, rather than as context that either helps or hurts on its own. For research, it connects two lines of work that usually stay separate: recurrent and memory-based reuse of past key-value pairs (Transformer-XL, Memorizing Transformers, StreamingLLM) and cache-reuse methods for repeated text (Prompt Cache, CacheBlend). It also offers a mechanistic account of why retained history has inconsistent effects, showing that a small, problem-specific component of representation change is stable across histories. The joint LoRA result suggests the approach complements weight adaptation rather than replacing it.
Real-world applications:
- Multi-turn assistants where a user works through many independent problems in a single session, such as homework help or exam preparation, and where earlier turns currently degrade later answers on mathematical reasoning.
- Cost-conscious deployment of frozen or barely tunable models, since only 12,288 parameters are trained and the backbone, stored keys and values, and output projection all stay fixed.
- Conversation-serving infrastructure that must decide how much session history to keep and how to read it, given that limiting the bank from the full set to the first 8K tokens drops T2 Avg@4 from 48.33% to 43.33%.
- Scientific and technical question answering is a partial fit: the Qwen3-4B Instruct model improved on GPQA-Diamond, but the Qwen3.5 models did not, and the authors flag this domain gap explicitly.
Industry relevance centers on the trade-off the paper quantifies: gains of up to 11.67 percentage points in mean T2-T4 Avg@4 over the unmodified model come with a median 0.81 s and 6.60 GiB peak memory overhead for full-bank reading, and bank storage grows as completed turns accumulate. That makes the method most attractive where accuracy on later turns matters more than prefill latency, and it points to selective retention as the natural lever for production settings.
Future Directions
- Broader training mixtures. Reflector directions are fitted only on mathematical problems, and the authors suggest that these directions may adapt to query-history relationships frequent in the training distribution. Testing other mixtures would clarify whether this contributes to the gap between AIME 2025 and GPQA-Diamond.
- Longer sessions and selective retention. Longer conversations increase bank storage and prefill reading costs as each completed turn adds new states; the authors propose selective retention to limit bank growth, and the fixed-length runtime measurements give a baseline for that work.
- Pretraining reflector directions jointly with the backbone. The authors suggest that pretraining on diverse task switches could learn these directions alongside the model, and that supervised fine-tuning or reinforcement learning could test how the readout adapts as reasoning behavior changes.
- Consolidating the evidence base. Current evidence covers four-turn sessions within the Qwen family, so the generality of the effect across architectures, session lengths, and task types remains open.
Target Audience
Researchers working on LLM reasoning and multi-turn conversation, particularly those interested in attention-level interpretability, key-value cache reuse, and parameter-efficient adaptation. The paper is also useful for engineers deploying multi-turn conversational assistants who need to weigh later-turn accuracy against prefill latency and memory when deciding how much session history to retain. Readers who want only the high-level claim can read the introduction, Table 2, and the conclusion; the appendices assume familiarity with attention mechanics and representation geometry.
Authors’ abstract
Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.