Research
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Overview Research area: Mechanistic interpretability and parameter-efficient adaptation of transformer language models — specifically, what limits how far a pretrained model can follow chains of in-co

- arXiv
- 2609.36585
- Published
- 2026-09-29
- Authors
- Zehao Jin, Ruixuan Deng, Junran Wang
AI summary
Overview
Research area: Mechanistic interpretability and parameter-efficient adaptation of transformer language models — specifically, what limits how far a pretrained model can follow chains of in-context references, and how small a change is needed to remove that limit.
Technical level: Advanced. The paper assumes familiarity with residual-stream activations, LoRA, causal tracing, attention knockouts, and linear probing, though its central claims are stated in plain terms.
Scope: The paper shows that thirteen pretrained base models and three looped models follow only very short reference chains by default, traces the computation that stops early, and demonstrates that a rank-8 LoRA at a single early layer starts a "relay" that frozen middle layers carry much further.
What This Paper Is About
A transformer can answer a prompt like K = apple; B = K; D = B; print(D) without any outside knowledge, but adding a few more assignments breaks it. The authors measure exactly how short this default computation is across thirteen standard base models, then show that training a rank-8 LoRA at one layer — while keeping every pretrained weight frozen — extends chain-following dramatically. The paper's second half explains where that computation lives inside the network and where an intervention must be placed to work.
Key Contributions
-
A short default computation, measured across models. Thirteen standard base models from Qwen3, Llama, OLMo-3, and Gemma-3, plus DeepSeek-V4-Flash (292B MoE), reliably follow only 1.4–3.6 lines (median 2.2). Extra pretrained loops in Ouro-1.4B, Ouro-2.6B, and Huginn-0125 add little.
-
A tiny edit with large gains. A rank-8 LoRA acting on the residual stream at one layer — 65,537 added parameters in Qwen3-8B, under 0.01% of the model — raises exact accuracy on 24-line chains from 15.5% to 99%. A longer-trained version reaches 50 lines in one pass, and in Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight.
-
A causal account of the relay. The LoRA causes program lines to pass on their chain identity through a short range of middle layers; frozen attention heads read progressively further up each chain, and removing parent-line attention stops the relay. Attention, probing, and causal restoration all support the same picture.
-
A prospective placement test, applied to MuSiQue. A measurement on the frozen model (the "cutoff layer") predicts where an intervention will stop working, meeting the preregistered criterion on three of four held-out models. Task-specific LoRAs on MuSiQue show the same preference for early layers.
Main Findings
- Default reach is short and depth does not fix it. Reliable reach across the thirteen standard models is 1.4–3.6 lines, median 2.2. Qwen3-8B drops from 83.5% choice accuracy at three lines to 54% at four and 41.5% at five. OLMo-3-32B and OLMo-3-7B both reach about 2.6 lines despite having 64 and 32 layers. DeepSeek-V4-F
Authors’ abstract
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/