Papers
arxiv:2609.36585

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Published on Sep 29
· Submitted by
Zehao Jin
on Oct 2
Authors:
,
,

Abstract

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/

Community

Paper submitter

TL;DR: Asked to answer directly (no CoT), LLMs follow reference chains like K = apple; B = K; D = B; print(D) for only a few lines. A tiny rank-8 LoRA at one early layer lets the frozen middle layers carry the chain much further. The computation is there; it just stops too early.

Key findings

  • 13 base models (0.6B–32B) reliably follow only 1.4–3.6 lines. Doubling depth (OLMo-3 7B → 32B) leaves reach at ~2.6 lines.
  • Qwen3-8B + a task-trained rank-8 LoRA at layer 14 (65,537 params, all base weights frozen): exact accuracy on 24-line chains goes from 15.5% → 99%. A longer-trained version reaches 50 lines in a single forward pass.
  • Mechanism: the LoRA acts on each token independently and moves no information between tokens. It starts a relay: program lines pass their chain identity along through frozen layers 16–22. Cutting each line's attention to its parent in layers 14–22 drops accuracy to chance; the same cut in layers 23–29 barely matters.
  • Placement cliff: same recipe, LoRA at layer 20 → 20.5 lines; at layer 21 → 5.2 lines. A frozen-model measurement located this limit within a preregistered tolerance in 3 of 4 held-out models.
  • Looped models: in Ouro-1.4B, a LoRA applied every loop reaches 60 lines after 4 loops and ≥160 after 8 (two-chain choice accuracy).
  • Multi-hop QA: on MuSiQue (gold paragraphs), separately trained early-layer LoRAs add 9.4–17.9 EM across three standard models.

🎬 Website : https://lunamos.github.io/stop-thinking-too-early/
💻 Code: https://github.com/Lunamos/stop-thinking-too-early

Happy to answer questions!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36585
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36585 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36585 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36585 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.