Skip to content
AI.info

Research

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Overview Research area: Mechanistic interpretability and parameter-efficient adaptation of transformer language models — specifically, what limits how far a pretrained model can follow chains of in-co

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
arXiv
2609.36585
Published
2026-09-29
Authors
Zehao Jin, Ruixuan Deng, Junran Wang

AI summary

Overview

Research area: Mechanistic interpretability and parameter-efficient adaptation of transformer language models — specifically, what limits how far a pretrained model can follow chains of in-context references, and how small a change is needed to remove that limit.

Technical level: Advanced. The paper assumes familiarity with residual-stream activations, LoRA, causal tracing, attention knockouts, and linear probing, though its central claims are stated in plain terms.

Scope: The paper shows that thirteen pretrained base models and three looped models follow only very short reference chains by default, traces the computation that stops early, and demonstrates that a rank-8 LoRA at a single early layer starts a "relay" that frozen middle layers carry much further.

What This Paper Is About

A transformer can answer a prompt like K = apple; B = K; D = B; print(D) without any outside knowledge, but adding a few more assignments breaks it. The authors measure exactly how short this default computation is across thirteen standard base models, then show that training a rank-8 LoRA at one layer — while keeping every pretrained weight frozen — extends chain-following dramatically. The paper's second half explains where that computation lives inside the network and where an intervention must be placed to work.

Key Contributions

  1. A short default computation, measured across models. Thirteen standard base models from Qwen3, Llama, OLMo-3, and Gemma-3, plus DeepSeek-V4-Flash (292B MoE), reliably follow only 1.4–3.6 lines (median 2.2). Extra pretrained loops in Ouro-1.4B, Ouro-2.6B, and Huginn-0125 add little.

  2. A tiny edit with large gains. A rank-8 LoRA acting on the residual stream at one layer — 65,537 added parameters in Qwen3-8B, under 0.01% of the model — raises exact accuracy on 24-line chains from 15.5% to 99%. A longer-trained version reaches 50 lines in one pass, and in Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight.

  3. A causal account of the relay. The LoRA causes program lines to pass on their chain identity through a short range of middle layers; frozen attention heads read progressively further up each chain, and removing parent-line attention stops the relay. Attention, probing, and causal restoration all support the same picture.

  4. A prospective placement test, applied to MuSiQue. A measurement on the frozen model (the "cutoff layer") predicts where an intervention will stop working, meeting the preregistered criterion on three of four held-out models. Task-specific LoRAs on MuSiQue show the same preference for early layers.

Main Findings

  • Default reach is short and depth does not fix it. Reliable reach across the thirteen standard models is 1.4–3.6 lines, median 2.2. Qwen3-8B drops from 83.5% choice accuracy at three lines to 54% at four and 41.5% at five. OLMo-3-32B and OLMo-3-7B both reach about 2.6 lines despite having 64 and 32 layers. DeepSeek-V4-F

Authors’ abstract

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/

Read the original paper