Training a Looped Transformer: When to Repeat, Review, and Stop
I started with a fairly innocent idea: let a small transformer reuse the same block several times, and see what it could learn through repeated computation.
Then it learned. Then it learned something else. Then the first lesson became much harder to retrieve.
That is where the interesting part started.
I chose a looped transformer because it lets me vary the number of internal passes while reusing the same block parameters. The experiments begin by testing that choice, then follow the training problem it exposed: how to retain earlier lessons while learning later ones. The retention methods are potentially useful beyond this architecture; recurrence itself still has to earn its place.
My instinct as an instructor was to try a refresher, check what had slipped, and avoid hammering the same lesson until something else broke. Human teaching supplied the questions. Controlled model experiments had to supply the answers.
My explanations often start as pictures. Sometimes they are crayon drawings: a filing cabinet, a pendulum, a student who needs a reminder. The picture helps me decide what to test. The measurements tell me which parts of the picture I get to keep.
This is an early experimental note from the Looped Long-Horizon Learner (LLHL) research project. Its aim is to investigate how a model can acquire new abilities over a sequence of tasks, preserve earlier ones, and use targeted review when performance slips. “Long-horizon” refers here to that extended learning history. This article examines the training method through a small, finite symbolic task whose answers can all be checked against an exact reference calculation.
The results include a useful recovery effect, a successful review policy, and a couple of ideas that looked better before we measured them properly.
Two loops, two different jobs
A looped transformer (LT) reuses a transformer block across depth: the same processing module runs several times within one prediction. Its hidden state—the intermediate numerical representation of the input—changes at each pass while the block's learned parameters are shared. Recurrent computation of this kind has established precedents, including Universal Transformers.
The distinction can be drawn without an expensive diagram:
Untied depth: input → block A → block B → block C → output
Tied depth: input → block W → block W → block W → output
same parameters, evolving state
For this model, the inner computation is approximately:
Here, is the hidden state at pass , is the shared block with learned parameters , and identifies the loop step. The final hidden state goes through a normalization layer and a prediction head, which produces a score for each possible answer.
The inner loop updates activations. It does not update the weights each time it runs. Ordinary backpropagation then accumulates contributions through the repeated uses of the shared block, and the optimizer updates the parameters.
The outer training loop decides which examples to teach next, what to review, how strongly to weight those objectives, and when to stop. I use curriculum for that ordering and mixture of lessons, and refresher or review for additional training on previously taught examples.
Those are separate design decisions. More internal passes do not automatically give a model better retention between lessons.
What changes compared with a conventional transformer setup?
| Decision | Conventional untied stack | This looped setup |
|---|---|---|
| Parameters across depth | Each block has its own parameters | One block is reused |
| Gradient contributions | Each block receives gradients for its role | Shared parameters receive contributions from every use |
| Increasing depth | Usually adds blocks and parameters | Adds passes without adding block parameters |
| Compute and memory | Depend on depth and implementation | Repeated passes still cost compute; training activations can still grow with depth |
| Evaluation | Usually uses the trained depth | Different loop counts can be tested, but improvement is not guaranteed |
The optimizer is still an ordinary optimizer. Our curriculum and refresher experiments could also be applied to other architectures. This study does not establish that their benefits are unique to looped transformers.
A small model with answers we can actually check
The trained module has 54,106 parameters, width 64, four attention heads, a feed-forward width of 256, GELU activation, pre-layer normalization, and no dropout. It learns a collection of fixed one-to-one symbol mappings: within each mapping, every input symbol has exactly one output, and each output is used once. For illustration, a three-symbol mapping could be a → c, b → a, c → b; the experiment uses larger mappings. Learned numerical vectors called embeddings identify which mapping is requested and which input position is being processed. Loop-step features tell the shared block which pass it is performing.
There are 260 labeled input–output cases across those mappings, split into two lessons of 130 each, which I will call A and B. The reported experiments directly supervise those cases with cross-entropy, a loss that penalizes giving the correct answer too little probability. They run on a central processing unit (CPU) using 32-bit floating-point numbers; they are not graphics processing unit (GPU) throughput benchmarks.
That deliberately narrow setup gives us exact answers and makes interference—learning one lesson harming performance on another—easy to count. It also sets a clear boundary: accuracy on these mappings measures acquisition and retention of the taught material. An initialization seed controls the random starting weights; matched seeds let competing methods start from the same model. New seeds test robustness across training runs, not generalization to previously unseen mappings.
We measure two outputs:
- Raw accuracy: the model's highest-scoring answer for each input.
- Projected accuracy: accuracy after an assignment algorithm chooses outputs jointly to enforce a valid one-to-one mapping. If two inputs both prefer the same output, this step resolves the conflict using the model's scores.
The projection uses SciPy's linear_sum_assignment(..., maximize=True) on the score matrix to maximize the total assigned score under the one-to-one constraint.
A structural constraint can repair an invalid set of choices without fixing the underlying scores. Later experiments require both measurements to pass. Otherwise, postprocessing can hide early regression.
More loops were not automatically more economical
The first temptation was obvious: if repeated computation helps, give the model more passes.
We compared one, two, four, and eight passes from matched initializations, using three seeds, AdamW, a learning rate of 0.001, and 100 optimizer updates per branch. AdamW is the rule used to adjust the learned parameters; the learning rate controls the scale of those adjustments. A branch is a separate training run. Supervision was applied after the final pass. Evaluation occurred every ten updates.
| Loops per forward pass | First observed perfect projected fit: three seeds | Block passes at those checkpoints |
|---|---|---|
| 1 | 70 / 70 / 70 updates | 70 / 70 / 70 |
| 2 | 60 / 60 / 70 updates | 120 / 120 / 140 |
| 4 | 50 / 50 / 60 updates | 200 / 200 / 240 |
| 8 | 50 / 50 / 60 updates | 400 / 400 / 480 |
All twelve branches eventually fit the projected mappings perfectly. More loops sometimes reduced the number of optimizer updates, but increased the number of block passes spent reaching that checkpoint.
At a common budget of 80 block passes, mean projected accuracy was 260/260 for one loop, 125.3/260 for two, 58.7/260 for four, and 37.0/260 for eight.
At that budget, the eight-loop model had received only ten optimizer updates, compared with eighty for the one-loop model. This measures progress under a shared block-pass budget; it does not mean the eight-loop model lost an ability it had already acquired.
Three matched seeds. “Block passes” means loop count multiplied by optimizer updates; it is a compute proxy, not an exact count of arithmetic operations or a measurement of elapsed time. First-perfect checkpoints were observed only every ten updates.
For this task and recipe, one pass was the economical reference. We had not demonstrated an advantage from recurrence. A broader depth-by-learning-rate search could change the ranking: the learning rate used here had been selected at depth two.
Increasing inference depth was also not monotonically helpful. In an earlier short screen, one model trained with two loops scored 64/260 at two and four inference loops, then fell to 53/260 at eight. Extra passes need their own evaluation.
The cool factor survived. The assumption that more loops meant a better bargain did not.
Learning again was faster than learning from scratch
Next came a different question: after a model has learned A and then lost much of its A performance while practicing B, what does it take to recover?
The later recovery and review comparisons held the architecture at two loops, using AdamW at 0.003 except where learning rate was the explicit treatment. Holding that recipe fixed let us compare curricula; it did not make two loops the compute-optimal choice.
A checkpoint is a saved training state. It can hold the learned weights along with the optimizer's running statistics and other information needed to resume training. We compared four starting histories, matched by initialization seed:
- A for 100 updates, followed by B for 100 updates.
- B alone for 200 updates, matching the total training lifetime.
- B alone for 100 updates, matching B exposure.
- The original untrained checkpoint.
Every branch received a fresh optimizer before learning or reviewing A. That removes inherited optimizer momentum as the explanation for faster recovery, while deliberately preserving the different weight histories.
Mastery meant five consecutive perfect observations on both raw and projected A outputs. The table records the first observation in that qualifying streak; it was confirmed four updates later.
| Starting history | Updates to A mastery, three seeds |
|---|---|
| Learned A, then practiced B | 7 / 7 / 5 |
| B only, 200 prior updates | 18 / 16 / 18 |
| B only, 100 prior updates | 19 / 16 / 18 |
| Fresh initialization | 27 / 22 / 27 |
Each dot is one of three matched seeds. The result concerns optimizer updates, not elapsed-time speedup.
The previously learned skill recovered with 75% fewer updates on average than fresh acquisition, and 63.5% fewer than the equal-lifetime B-only control.
That supports a recovery advantage associated with prior learning. It does not tell us whether recovery restored an earlier computational route, repaired a partly preserved representation, or built another useful route. The impaired models still answered roughly 30 of the 130 A cases correctly before review, so this was recovery after substantial impairment, not recovery from complete observable loss.
My picture was a filing cabinet: perhaps some of the material is still there, but the learner has lost track of where it was filed. A refresher might help it find the material faster than rebuilding the whole collection.
The faster recovery makes that picture worth investigating. We have not demonstrated that the contents survived intact or that a missing index was the cause. The model does not come with labeled drawers we can open. Distinguishing preserved content from reconstructed computation takes another experiment.
Repairing A could damage B
There was a catch. Continuing to train only A eventually harmed B.
At the moment A first recovered, the three models still got 125, 118, and 127 of the 130 B cases right in their raw outputs. After 60 A-only updates, those counts had fallen to 73, 64, and 66.
Projected outputs initially concealed some of that damage. This is why measuring only the repaired skill, or only the constrained output, would have given us an incomplete account.
The next experiment used what I call a petri rack: a controlled collection of separate training branches, like cultures grown under different conditions. Each branch starts from a preserved parent checkpoint, the common saved starting state for that comparison, with explicit changes to learning rate and lesson weights. Coarse searches narrowed into smaller searches, always returning to the same parent for each seed. All branches remained archived, including runs stopped at their update limit and unsuccessful recipes.
The chosen recipe restored A while keeping B perfect at every measured update on three fresh validation seeds—new starting models reserved for checking the selected recipe. It used learning rate 0.003 and A:B loss weights of 1:2, giving B's loss twice A's weight in the training objective. It stopped after 15 updates in each seed, under a rule requiring five consecutive observations with every A and B answer correct in both raw and projected outputs, and stopping checks at five-update boundaries.
Balanced 1:1 review was faster in two seeds and briefly lost one B answer in one seed. The selected recipe won because our stated selection rule prioritized preservation over speed. It was a trade-off we chose, not a universal optimum the model discovered.
Rehearsal—revisiting earlier examples while learning new material—is established practice in continual learning, where a model learns through a sequence of tasks. Experience replay is one relevant precedent. Our practical question was how to choose and stop the refresher in this particular controlled setting.
“Ease up as the gap closes” needed a control
Another intuition was to reduce training pressure as the errors disappeared.
I pictured a pendulum with controllable magnets. Pull hard enough to bring it back toward the desired position, then ease off as it approaches. Keep pulling at full strength and you might send it too far the other way.
In the model experiment, the measurable controls were learning rate, loss mixture, and stopping. The pendulum suggested a feedback rule; it did not establish that the optimization dynamics actually behaved like a pendulum.
We tested a fixed learning rate, a preset taper that reduced it on a schedule, and a taper driven by the remaining error count. The feedback rule kept a nonzero learning-rate floor so corrective learning could continue near mastery.
On three fresh validation seeds:
| Policy | Stopping updates | Mean parameter-path reduction versus fixed |
|---|---|---|
| Fixed rate | 15 / 15 / 10 | Reference |
| Preset taper | 15 / 15 / 10 | 21.5% |
| Error-gap taper | 15 / 15 / 15 | 24.1% |
“Parameter path” here means the sum of the lengths of the actual parameter changes at each update, measured with the L2, or Euclidean, norm. Think of total distance traveled through weight space, including any backtracking. It is neither computation saved nor a measure of damage avoided.
The feedback taper moved the weights less but took more updates on average. It failed our declared adoption rule, which required no update-count penalty versus the fixed control.
Then we tested the recovered models under another imbalanced practice challenge: practice A and measure errors on B, or practice B and measure errors on A. The fixed-rate parents accumulated less error on the lesson left out of practice than either tapered parent in all six comparisons: three seeds, each tested in both directions. All challenged models could subsequently be repaired, but less earlier movement had not bought better retention.
That result matters. “Gentler” sounded sensible. We still needed to measure what it protected.
A refresher can start before an answer becomes wrong
Accuracy is binary. The underlying score difference is not.
For a labeled example, define its margin as:
Here, is the model's score for answer , before converting scores into probabilities. The margin is the correct answer's score minus its strongest competitor's score. A correct answer can retain the lead while that lead becomes much smaller. Our next review policy used the remaining fraction of the margin measured at a preserved, fully correct parent checkpoint. For example, a margin falling from 4 to 1 has a current-to-parent ratio of 0.25: the answer is still correct, but its lead has shrunk to one quarter of its earlier size.
It selected:
- Existing errors first.
- Then correct examples with the lowest current-to-parent margin ratio.
The parent margins were positive. This ratio is a diagnostic score, not a calibrated probability of forgetting.
Written explicitly, the ratio is for example . Correctness is determined by the model's actual highest-scoring prediction: a zero margin is a tie and is not automatically a wrong answer. The implemented selector orders wrong answers by their current margin, most negative first, then fills the remaining places with correct answers in ascending ratio order. Ties among those ratios are broken by current margin, then by a stable example index; errors with equal margins also use that index.
The useful contrast here is with accuracy: an answer can remain correct while its margin shrinks. Cross-entropy loss can also change before an answer becomes wrong. We have not compared this policy with review selected by loss, so these results do not establish that margins provide an earlier warning than loss.
We compared three policies with the same supervision budget: practice all 130 cases in one group, then review 26 cases from the other. Each selected case contributed equally to the mean loss. A no-review reference received no extra review cases and therefore was not an equal-budget treatment.
The policy was fixed after a preliminary trial, then tested on three fresh parent seeds in two practice directions: practice A while reviewing B, and practice B while reviewing A. These are six related branches, not six independent initializations. Uniform rotating review cycles through the earlier lesson's examples in turn; error-first review prioritizes currently wrong answers and fills spare review places by rotation. The margin policy instead fills those spare places with the correct answers whose margins have weakened most relative to the parent.
| Policy | Mean cumulative review-group errors over 40 updates | Final perfect branches |
|---|---|---|
| No review | 729.83 | 0/6 |
| Uniform rotating review | 33.17 | 6/6 |
| Error-first review | 9.33 | 6/6 |
| Errors, then weakening margins | 0.00 | 6/6 |
Cumulative errors count a case again each time it is wrong at a post-update evaluation. They are not counts of unique failed examples. The figure compares the three equal-supervision-budget review policies.
The margin policy preserved every raw and projected answer at all 40 measured updates in all six branches. The other review policies also finished perfect. Looking only at the final score would have hidden their temporary regressions.
There is a substantial information advantage built into this experiment: all 260 labels remained available for diagnostics, even when only 156 cases contributed to the loss. We computed all mappings, inspected the scores, and selected the review subset. Equal review budgets do not establish equal wall time, reduced label access, or less transformer computation.
Repeated lessons, with an important tie
When teaching people, I like to end the day with a recap and begin the next one by revisiting the previous lesson. Sometimes a learner needs the reminder that they already know how to do something. That observation suggested a sequence we could test with a model: practice, review, switch, review again.
A paused model does not consolidate memories in its sleep. Here, a review means actual labeled examples and actual parameter updates. We extended the comparison to six alternating A/B phases of 40 updates each. This produced 24 trajectories: three seeds, two starting orders, and four policies. Model weights, optimizer state, random state, and counters continued across phase boundaries.
Across all six phases, the margin policy remained perfect at every measured update. However, uniform and error-first review also accumulated zero errors during phases two through six.
The overall difference came from the first phase, which had already been tested. This extension demonstrates sustained retention for all three review policies in that sequence. It does not establish a new margin-policy advantage during the later switches, and it reused the fresh-confirmation parents rather than providing an independent replication.
Sometimes the useful result is that several methods keep working.
What the petri rack preserves
The reusable framework keeps the experiment's history explicit:
- An immutable parent for each matched comparison.
- Separate model, optimizer, random-number state, and counters for each branch.
- Recorded recipes, source snapshots, metrics, and checkpoint hashes.
- A declared distinction between exact continuation (resume all saved training state), a changed fork (branch from that state under a new recipe), and a weights-only restart (keep learned weights but start a fresh optimizer).
- Selection criteria fixed before fresh-seed validation.
- Raw predictions and constrained outputs evaluated separately.
In the repeated-lesson extension, the recorded audit replayed all 5,760 review decisions from pre-update checkpoints, checked 120 complete phase transitions, and independently scored 144 phase endpoints. Those checks establish what the experiment did. They do not turn three seeds into a large population study.
The framework supports independent branches. In its joint execution mode, several models' calculations are combined for an optimizer step while their parameters remain separate. This has not established that a rack of distinct models can train efficiently together on a GPU. That performance question remains open.
What I would carry into the next experiment
The useful changes to my training procedure are concrete:
- Count internal passes as well as optimizer updates. A shorter training curve can conceal more computation.
- Test inference depth separately. Extra loops can make an answer worse.
- Measure retention while training continues. A final perfect score can hide the route taken to get there.
- Compare recovery with fresh acquisition and matched training-history controls. Faster repair needs an appropriate reference.
- Select refresher content from measured gaps. Wrong answers and weakening correct answers can carry different information.
- Stop under a declared all-skill rule. Exclusive remediation can outlive its usefulness.
- Keep the losing branches. They are often where the explanation becomes more honest.
These are results from a small directly supervised task. We still need unseen task families, stronger replay-selection baselines, comparable untied models, measured compute budgets, and tests of genuinely new capabilities. Additional controls must also establish whether any benefit comes specifically from recurrence.
The longer-term LLHL question is whether a learner can keep accumulating useful abilities while making review selective, measurable, and affordable. These experiments are a small step toward asking that question properly.
My shorthand is the R2-D2 factor. In Star Wars, the resourceful little droid retains memories that routine wipes would have erased; Anakin breaks protocol by leaving his memory intact. The idea I borrow is that a familiar machine can have a richer history to draw on when improvising because it carries its experience from one assignment to the next.
A preserved checkpoint makes that experiment possible for our models; it does not guarantee useful accumulated experience. We still have to measure what transfers, what interferes, and what needs a refresher.
They also connect with the question in my earlier Same Score, Different Memory experiment: what does a perfect score fail to tell us about the next lesson? That work used a different model and intervention; this is a related question, not a replication.
For now, the model can learn the task, suffer interference, recover, and sometimes avoid measurable regression through a better-chosen refresher. The mechanism of that recovery remains open.
That is what I mean by crayon doodling in 8K: keep the simple picture, then increase the resolution of the test.
I can still joke that it is magnets. The measurements have to be more specific.



