Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Abstract
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
Community
TL;DR: Asked to answer directly (no CoT), LLMs follow reference chains like K = apple; B = K; D = B; print(D) for only a few lines. A tiny rank-8 LoRA at one early layer lets the frozen middle layers carry the chain much further. The computation is there; it just stops too early.
Key findings
- 13 base models (0.6B–32B) reliably follow only 1.4–3.6 lines. Doubling depth (OLMo-3 7B → 32B) leaves reach at ~2.6 lines.
- Qwen3-8B + a task-trained rank-8 LoRA at layer 14 (65,537 params, all base weights frozen): exact accuracy on 24-line chains goes from 15.5% → 99%. A longer-trained version reaches 50 lines in a single forward pass.
- Mechanism: the LoRA acts on each token independently and moves no information between tokens. It starts a relay: program lines pass their chain identity along through frozen layers 16–22. Cutting each line's attention to its parent in layers 14–22 drops accuracy to chance; the same cut in layers 23–29 barely matters.
- Placement cliff: same recipe, LoRA at layer 20 → 20.5 lines; at layer 21 → 5.2 lines. A frozen-model measurement located this limit within a preregistered tolerance in 3 of 4 held-out models.
- Looped models: in Ouro-1.4B, a LoRA applied every loop reaches 60 lines after 4 loops and ≥160 after 8 (two-chain choice accuracy).
- Multi-hop QA: on MuSiQue (gold paragraphs), separately trained early-layer LoRAs add 9.4–17.9 EM across three standard models.
🎬 Website : https://lunamos.github.io/stop-thinking-too-early/
💻 Code: https://github.com/Lunamos/stop-thinking-too-early
Happy to answer questions!
Get this paper in your agent:
hf papers read 2609.36585 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper