Papers
arxiv:2610.11570

Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals

Published on Oct 8
· Submitted by
Pengxiang Li
on Oct 9
Authors:
,
,
,
,
,

Abstract

In this paper, we argue that looped Transformers need their own residual connections to prevent performance degradation as the number of iterations grows. We observe that increasing loop iterations can reduce reasoning accuracy: noisy state updates overwrite correct intermediate deductions and even undo completed solutions. This leaves subsequent iterations to recover lost information from an already degraded representation: once an error arises in an earlier loop, often as a result of long-range propagation through the recurrence, later loops find it difficult to correct. In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update. InfiLoop combines content-based weighting with learned temporal decay to maintain a running summary of recurrent states. An exact streaming recurrence keeps its persistent aggregation memory constant as the loop count grows. The resulting adaptive update suppresses unreliable proposals and preserves useful intermediate states. Across extensive reasoning tasks, a 7M-parameter InfiLoop model outperforms existing recursive architectures, reaching 97.9% exact accuracy on Sudoku-Extreme, and 13.6% pass@2 on ARC-AGI-2. Notably, on Sudoku-Extreme, InfiLoop continues to improve with test-time looping beyond 20,000 effective steps, showing that added depth translates directly into stronger reasoning. Our code is available at https://github.com/pixeli99/InfiLoop.

Community

Paper submitter

1/ More loops ≠ better reasoning. On Sudoku-Extreme, TRM peaks at 62% and then drops as you keep looping. Puzzles it already solved get broken again.

2/ The culprit is “carry-last”: every iteration just overwrites the state with the latest output. One noisy step and the good answer is gone.

3/ InfiLoop keeps a running, attention-weighted average of past states instead. Each new update gets a learned step size. Good updates move the state; noisy ones barely touch it. Memory stays constant no matter how many loops you run.

4/ Results with 7M params: 97.9% on Sudoku-Extreme, 13.6% pass@2 on ARC-AGI-2, and it’s still improving at 20k+ effective steps. Not best everywhere (Maze-Hard is about on par with prior work), but the depth scaling is the point.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.11570 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.11570 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.11570 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.