File size: 2,871 Bytes
36295d9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
---
license: mit
library_name: pytorch
pipeline_tag: text-generation
tags:
- nested-learning
- hope
- test-time-learning
- test-time-training
- continual-learning
- byte-level
- attention-free
- recurrent
- arxiv:2512.24695
---

# HOPE — Nested Learning in PyTorch

The index card for [nested-learning-hope](https://github.com/smallhours19/nested-learning-hope),
an independent PyTorch reproduction of HOPE from **"Nested Learning: The Illusion
of Deep Learning Architectures"** ([arXiv:2512.24695](https://arxiv.org/abs/2512.24695),
Behrouz et al., Google Research). No official code or weights exist for the paper;
this is an unofficial reproduction at reduced scale.

HOPE is an attention-free language model whose weights change during the forward
pass (test-time learning): Self-Modifying Titans (Eq. 83-93, gated delta rule,
self-generated targets) plus a Continuum Memory System (Eq. 70-71, MLP levels
updated every 1/2/4 chunks by the accumulated task-loss gradient).

## The released recipe (reproducible in ~7 GPU-hours)

| | |
|---|---|
| Parameters | 30.6M (4 layers, d=512, chunk 512, no positional embedding) |
| Tokenizer | none — raw bytes (vocab 256) |
| Training data | FineWeb-Edu sample, 0.3B bytes, sequence length 2048 |
| Held-out validation | **1.721 bits/byte** |
| Needle recall (1K / 6.6K bytes) | **+2.59 / +2.10 bpb** vs control (n=40) |
| 8K extrapolation (trained at 2K) | flat, no horizon (1.66 → 1.96 bpb) |

Every number is reproducible from the GitHub repository (committed training
log, evaluation scripts, and an exactness test verifying the chunk-parallel
path equals the sequential recurrence to fp64 machine precision).

## Reproducing the model

No weights are hosted here (yet) — the checkpoint reproduces from the GitHub
repository in about 7 hours on one 16 GB GPU, and the committed training log
lets you verify your run against ours at step 100/200 before committing the
full budget:

```bash
git clone https://github.com/smallhours19/nested-learning-hope && cd nested-learning-hope
pip install -r requirements.txt
python scripts/download_data.py --out data/fineweb-edu --shards 3
python scripts/train_byte.py --data data/fineweb-edu --out runs/hope-byte
```

Where this implementation had to choose beyond the paper's text, every choice
is documented in
[DEVIATIONS.md](https://github.com/smallhours19/nested-learning-hope/blob/main/docs/DEVIATIONS.md);
stability devices the paper does not mention are cataloged in
[STABILITY.md](https://github.com/smallhours19/nested-learning-hope/blob/main/docs/STABILITY.md).

## Citation

```bibtex
@article{behrouz2025nested,
  title   = {Nested Learning: The Illusion of Deep Learning Architectures},
  author  = {Behrouz, Ali and Razaviyayn, Meisam and Zhong, Peilin and
             Mirrokni, Vahab},
  journal = {arXiv preprint arXiv:2512.24695},
  year    = {2025}
}
```