File size: 19,240 Bytes
fee196d
8398a43
 
fee196d
8398a43
 
 
 
 
 
 
 
 
fee196d
 
5f38bca
eb753bc
8398a43
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
---
language: en
license: apache-2.0
tags:
- causal-lm
- research
- fp8
- attention
- normalization
- neollm
- pace
datasets:
- HuggingFaceFW/fineweb-edu
---

# NeoLLM

NeoLLM is a **135 M parameter** decoder-only language model trained from scratch on
[FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) in **FP8**
precision, completing training in approximately **6 hours** on a single NVIDIA RTX 5090.
It integrates a collection of recently published attention and normalization techniques
into a single architecture, with the goal of studying how they interact during
pretraining. The model is actively being developed and the current checkpoint represents
an intermediate training state.

> **Author / contact:** [@Kyokopom](https://x.com/Kyokopom) on X
> **Repository:** [KitsuVp/NeoLLM](https://huggingface.co/KitsuVp/NeoLLM)

---

## Architecture

NeoLLM is a decoder-only transformer with the following configuration:

| Parameter | Value |
|---|---|
| Hidden size | 512 |
| Layers | 12 |
| Attention heads | 8 |
| KV heads (GQA) | 4 |
| Head dim | 64 |
| Intermediate size | 1536 |
| Vocabulary | Qwen3 tokenizer (64,402 tokens) |
| Context length | 512 tokens |

### Parameter breakdown

| Parameter bucket | Count |
|---|---|
| **Total parameters** | 84.57M (84,569,432) |
| **Embedding parameters** (tied) | 32.97M (32,973,824) |
| **Non-embedding parameters** | 51.60M (51,595,608) |
| **Effective trainable parameters** | 84.57M (84,569,432) |

> Weight tying is **enabled**: the input embedding matrix and the language-model head
> share the same parameters, so the effective trainable budget is
> `total − embed = 51.60M`.

### Integrated techniques

NeoLLM combines architecture modules, optional auxiliary objectives, and
training-time optimizer/stability components from the following papers.

**Embedding and token representation**

- **Learnable Multipliers** ([arXiv:2601.04890](https://arxiv.org/abs/2601.04890)) — Adds
  per-row and per-column learnable scalar parameters to selected matrix layers and, when
  enabled, embeddings.
- **Leviathan** ([arXiv:2601.22040](https://arxiv.org/abs/2601.22040)) — Optional
  continuous token embedding generator that can replace the discrete input lookup table.
- **KHRONOS** ([arXiv:2505.13315](https://arxiv.org/abs/2505.13315)) — Kernel/basis
  reference used by the Leviathan continuous token generator implementation.
- **Spelling Bee Embeddings** ([arXiv:2601.18030](https://arxiv.org/abs/2601.18030)) —
  Augments token embeddings with character-level spelling information.
- **Token Embedding Manifold analysis** ([arXiv:2504.01002](https://arxiv.org/abs/2504.01002)) —
  Reference motivation for treating token embeddings as structured objects rather than
  unconstrained lookup rows.

**Attention, positions, and output projection**

- **FAN** ([arXiv:2502.21309](https://arxiv.org/abs/2502.21309)) — Fourier Analysis Networks.
  A portion of the projection channels are dedicated to periodic cosine/sine features.
- **MEA** ([arXiv:2601.19611](https://arxiv.org/abs/2601.19611)) — Explicit Multi-head
  Attention. Adds small learnable interaction matrices between attention heads for K and V.
- **LUCID** ([arXiv:2602.10410](https://arxiv.org/abs/2602.10410)) — Applies a learned
  lower-triangular preconditioner to V before attention, decorrelating value representations
  across positions.
- **Affine-Scaled Attention** ([arXiv:2602.23057](https://arxiv.org/abs/2602.23057)) — Adds
  two learnable per-head scalars (α and β) to the softmax weights:
  `[α·softmax(QKᵀ) + β]·V`.
- **XSA** ([arXiv:2603.09078](https://arxiv.org/abs/2603.09078)) — Exclusive Self Attention.
  After computing attention, removes the component of the output aligned with the token's
  own value vector.
- **Directional Routing** ([arXiv:2603.14923](https://arxiv.org/abs/2603.14923)) — Each head
  learns K=4 directions in the output space; a learned router suppresses the attention output
  along each direction per input.
- **Gated Attention** ([arXiv:2505.06708](https://arxiv.org/abs/2505.06708)) — A sigmoid gate
  is applied to the attention output before the output projection, introducing non-linearity
  and preventing attention sinks.
- **Momentum Attention** ([arXiv:2411.03884](https://arxiv.org/abs/2411.03884)) — Modifies Q
  and K by subtracting a fraction of the previous position's Q and K values (causal
  first-difference).
- **Interleaved Head Attention / IHA** ([arXiv:2602.21371](https://arxiv.org/abs/2602.21371)) —
  Builds pseudo-heads from learned cross-head mixtures to create multiple attention patterns
  per original head.
- **REPO** ([arXiv:2512.14391](https://arxiv.org/abs/2512.14391)) — Context re-positioning
  module that learns contextual position coordinates above a configurable start layer.
- **GRAPE** ([arXiv:2512.07805](https://arxiv.org/abs/2512.07805)) — Group representational
  position encoding used by the REPO-GRAPE positional path.
- **GOAT priors** ([arXiv:2601.15380](https://arxiv.org/abs/2601.15380)) — Optional
  factorized attention log-prior channels inspired by trainable attention priors.
- **Hadamard output projection** ([arXiv:2603.08343](https://arxiv.org/abs/2603.08343)) —
  Replaces dense attention output projection with a structured Hadamard transform plus
  lightweight scaling.

**Normalization, residual flow, and MLP**

- **SeeDNorm** ([arXiv:2510.22777](https://arxiv.org/abs/2510.22777)) — Applied to Q and K
  projections. Dynamically rescales normalization from the input's own statistics.
- **LayerNorm Scaling / LNS** ([arXiv:2502.05795](https://arxiv.org/abs/2502.05795)) — Each
  layer's output is scaled by 1/√ℓ where ℓ is the layer index.
- **GPAS** ([arXiv:2506.22049](https://arxiv.org/abs/2506.22049)) — Gradient-Preserving
  Activation Scaling for residual junctions.
- **PolyNorm** ([arXiv:2602.04902](https://arxiv.org/abs/2602.04902)) — Replaces the standard
  MLP activation with normalized linear, quadratic, and cubic branches.
- **SimpleGPT** ([arXiv:2602.01212](https://arxiv.org/abs/2602.01212)) — Second-order
  geometry-inspired normalization strategy applied inside MLP projections.
- **StackMemory / STACKTRANS** ([NeurIPS 2025](https://openreview.net/forum?id=2bbDg587uh)) —
  Optional differentiable hidden-state stack between decoder layers.
- **Attention Residuals / AttnRes** ([arXiv:2603.15031](https://arxiv.org/abs/2603.15031)) —
  Optional learned depth-wise aggregation over previous layer outputs or block summaries.
- **LAUREL** ([arXiv:2411.07501](https://arxiv.org/abs/2411.07501)) — Optional learned
  augmented residual layer with residual-weight and low-rank variants.

**Training objectives and training-time regularizers**

- **Cut Cross Entropy** ([Apple repository](https://github.com/apple/ml-cross-entropy)) —
  Memory-efficient next-token loss that avoids materializing the full token-by-vocabulary
  logits tensor. NeoLLM remains compatible with the upstream package when the extensions
  below are disabled.
- **MiLe Loss** ([arXiv:2310.19531](https://arxiv.org/abs/2310.19531)) — Optional detached,
  mean-normalized predictive-entropy weighting of token losses, implemented inside the
  extended CCE path.
- **Output Embedding Centering / mu-loss**
  ([arXiv:2601.02031](https://arxiv.org/abs/2601.02031)) — Optional
  `lambda * ||mean(output_embeddings)||^2` regularizer for output-logit stability.
- **MEAP** ([arXiv:2502.07490](https://arxiv.org/abs/2502.07490)) — Optional training-only
  input corruption that masks a fixed fraction of eligible tokens while preserving clean
  next-token labels, causal attention, and the inference path.
- **TWEO** ([arXiv:2511.23225](https://arxiv.org/abs/2511.23225)) — Optional
  Transformers Without Extreme Outliers activation regularizer for FP8/low-bit-friendly
  training.
- **NITP** ([arXiv:2605.24956](https://arxiv.org/abs/2605.24956)) — Optional Next Implicit
  Token Prediction auxiliary objective using shallow-layer implicit token targets and a
  cosine loss.
- **NextLat** ([arXiv:2511.05963](https://arxiv.org/abs/2511.05963)) — Optional next-latent
  prediction objective using latent dynamics, Smooth L1 supervision, and frozen-head KL.

### Optional extended-CCE configuration

| Feature | Enabled | Value |
|---|---:|---:|
| MiLe Loss | True | gamma=1.0 |
| mu-loss | True | lambda=0.0001 |
| MEAP | True | ratio=0.15 |

MiLe, mu-loss, and MEAP require the extended
[`Kitsunp/ml-cross-entropy`](https://github.com/Kitsunp/ml-cross-entropy) package only when
their corresponding flags are enabled. With all flags disabled, NeoLLM calls upstream CCE
without extension-specific arguments. When any extension is active, CCE reports three compact
scalars: unweighted NTP cross entropy, the MiLe reweighting delta, and the mu-loss penalty.
Their sum reconstructs `ntp_loss` exactly. MEAP reports its eligible and selected counts from
the masking kernel; the trainer logs the selected count and exact fraction. No diagnostic path
materializes full-vocabulary logits or a token mask outside the kernels.

**Optimizer and training stability**

- **Conda** ([arXiv:2509.24218](https://arxiv.org/abs/2509.24218)) —
  Column-Normalized Adam optimizer path used by the training script.
- **Cautious Weight Decay** ([arXiv:2510.12402](https://arxiv.org/abs/2510.12402)) —
  Sign-selective weight decay variant used by the custom optimizer logic.
- **Correction of Decoupled Weight Decay** ([arXiv:2512.08217](https://arxiv.org/abs/2512.08217)) —
  Adapts decoupled weight decay during learning-rate decay.
- **AdamHD** ([arXiv:2511.14721](https://arxiv.org/abs/2511.14721)) —
  Decoupled Huber decay regularization reference used by the optimizer.
- **GradientStabilizer** ([arXiv:2502.17055](https://arxiv.org/abs/2502.17055)) —
  Optional threshold-free gradient magnitude stabilizer.
- **PACE** ([arXiv:2606.25086](https://arxiv.org/abs/2606.25086)) —
  Optional iterate-average controller that trains for the EMA model returned at evaluation
  and final serialization. The Conda-basis adaptation and its difference from AdamW are
  documented below.

---

### PACE integration and AdamW-reference differences

PACE follows Au and Block's returned-model objective: the live weights are pulled toward a
power-law EMA with a clipped per-coordinate gain, and evaluation/final serialization use that
EMA estimator.

- **Reference AdamW rule:** the gain uses AdamW's original-coordinate diagonal
  second moment, `eta * c * (1+t)^(-kappa) / (sqrt(v_hat) + eps)`.
- **NeoLLM Conda rule (`mode=conda`):** for projected 2-D tensors, both the EMA
  displacement and `v_hat` are represented in Conda's cached SVD basis. The control is
  projected back after applying the diagonal gain. This is a deliberate change from AdamW
  required to avoid mixing incompatible coordinate systems.
- **Optional exact AdamW pullback geometry (`mode=adamw`):** an additional
  original-coordinate second moment is maintained for projected matrices. The live optimizer
  step remains Conda.
- **Conda scale:** in `mode=conda`, Conda's matrix-update scale multiplies the unsaturated
  gain automatically because it is part of the effective Conda preconditioner. AdamW has no
  corresponding scale.
- **Fixed algorithm internals:** the EMA is stored in FP32, the gain is clipped at `1`, and PACE
  reuses each Conda group’s numerical epsilon. These are not exposed as independent switches.
- **Minimal modes:** `use_pace=False` is plain Conda; `use_pace=True, c=0` is Conda+EMA;
  `use_pace=True, c>0` is complete PACE.
- **Ordering:** PACE runs only after Conda, CWD/CHD, and weight-decay correction have fully
  updated the live weights.
- **Disabled guarantee:** with `use_pace=False`, no PACE state is allocated and no existing
  Conda arithmetic or parameter update is changed.
- **Checkpoint policy:** resumable internal checkpoints retain live weights and complete optimizer
  state, while evaluation and the final returned/Hub model always use the EMA when PACE is active.

Current run: **disabled; no EMA state, auxiliary moment, or pullback is allocated**.

---

## Training

| Setting | Value |
|---|---|
| Dataset | FineWeb-Edu (sample-10BT) |
| Tokens seen | ~1.54B (46,875 steps × batch 64 × length 512) |
| Precision | FP8 native (E4M3 weights/activations, E5M2 gradients) + BF16 fallback |
| Optimizer | Conda (PACE disabled) |
| PACE | disabled; no EMA state, auxiliary moment, or pullback is allocated |
| Learning rate | 6e-04 with linear warmup (10 % of steps) |
| Weight decay | 0.1 |
| Training time | ~4h 14m |
| Hardware | NVIDIA RTX 5090 (single GPU) |

### Training curve

| Step | Train Loss | Val Loss |
|---|---|---|
| 5,000 | 5.898 | 5.449 |
| 10,000 | 5.829 | 5.364 |
| 15,000 | 5.682 | 5.204 |
| 20,000 | 6.148 | 5.526 |
| 25,000 | 5.551 | 5.051 |
| 30,000 | 5.410 | 4.911 |
| 35,000 | 5.560 | 5.120 |
| 40,000 | 5.364 | 4.864 |
| 45,000 | 5.246 | 4.722 |
| 46,875 | — | 4.671 |

---

## Limitations

- **Token budget** — ~1.5 B tokens seen; below estimated optimum. Knowledge-intensive tasks
  will improve with more training.
- **Gradient spike at step 40k** — Reorganized the attention pattern in layer 9 that
  previously captured long-range token correlations. A checkpoint from ~step 38k is expected
  to have better aggregate benchmark scores.
- **PolyNorm exclusivity** — The quadratic branch has become partially redundant with the
  linear branch. Will be corrected in the next training run.
- **Base model only** — Not instruction-tuned or aligned; purely a next-token-prediction
  base model.

---

## References

All papers whose techniques are integrated into NeoLLM's architecture,
training objective, or training stack:

| Area | Technique | Paper title | Reference |
|---|---|---|---|
| Embeddings | Learnable Multipliers | Freeing the Scale of Language Model Matrix Layers | [arXiv:2601.04890](https://arxiv.org/abs/2601.04890) |
| Embeddings | Leviathan | A Separable Architecture for Continuous Token Representation in Language Models | [arXiv:2601.22040](https://arxiv.org/abs/2601.22040) |
| Embeddings | KHRONOS | KHRONOS: a Kernel-Based Neural Architecture for Rapid, Resource-Efficient Scientific Computation | [arXiv:2505.13315](https://arxiv.org/abs/2505.13315) |
| Embeddings | Spelling Bee | Spelling Bee Embeddings for Language Modeling | [arXiv:2601.18030](https://arxiv.org/abs/2601.18030) |
| Embeddings | Token embedding analysis | Token Embeddings Violate the Manifold Hypothesis | [arXiv:2504.01002](https://arxiv.org/abs/2504.01002) |
| Attention / positions | FAN | Fourier Analysis Networks | [arXiv:2502.21309](https://arxiv.org/abs/2502.21309) |
| Attention / positions | MEA | Explicit Multi-head Attention for Inter-head Interaction in Large Language Models | [arXiv:2601.19611](https://arxiv.org/abs/2601.19611) |
| Attention / positions | LUCID | Attention with Preconditioned Representations | [arXiv:2602.10410](https://arxiv.org/abs/2602.10410) |
| Attention / positions | Affine-Scaled Attention | Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention | [arXiv:2602.23057](https://arxiv.org/abs/2602.23057) |
| Attention / positions | XSA | Exclusive Self Attention | [arXiv:2603.09078](https://arxiv.org/abs/2603.09078) |
| Attention / positions | Directional Routing | Directional Routing in Transformers | [arXiv:2603.14923](https://arxiv.org/abs/2603.14923) |
| Attention / positions | Gated Attention | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free | [arXiv:2505.06708](https://arxiv.org/abs/2505.06708) |
| Attention / positions | Momentum Attention | Momentum Attention | [arXiv:2411.03884](https://arxiv.org/abs/2411.03884) |
| Attention / positions | IHA | Interleaved Head Attention | [arXiv:2602.21371](https://arxiv.org/abs/2602.21371) |
| Attention / positions | REPO | Language Models with Context Re-Positioning | [arXiv:2512.14391](https://arxiv.org/abs/2512.14391) |
| Attention / positions | GRAPE | Group Representational Position Encoding | [arXiv:2512.07805](https://arxiv.org/abs/2512.07805) |
| Attention / positions | GOAT priors | You Need Better Attention Priors | [arXiv:2601.15380](https://arxiv.org/abs/2601.15380) |
| Attention / positions | Hadamard o_proj | Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers | [arXiv:2603.08343](https://arxiv.org/abs/2603.08343) |
| Residual / normalization | SeeDNorm | Self-Rescaled Dynamic Normalization | [arXiv:2510.22777](https://arxiv.org/abs/2510.22777) |
| Residual / normalization | LNS | The Curse of Depth in LLMs | [arXiv:2502.05795](https://arxiv.org/abs/2502.05795) |
| Residual / normalization | GPAS | Gradient-Preserving Activation Scaling | [arXiv:2506.22049](https://arxiv.org/abs/2506.22049) |
| Residual / normalization | PolyNorm | PolyNorm / PolyCom | [arXiv:2602.04902](https://arxiv.org/abs/2602.04902) |
| Residual / normalization | SimpleGPT | SimpleGPT | [arXiv:2602.01212](https://arxiv.org/abs/2602.01212) |
| Residual / normalization | StackMemory / STACKTRANS | Recursive Transformer: Boosting Reasoning Ability with State Stack | [NeurIPS 2025](https://openreview.net/forum?id=2bbDg587uh) |
| Residual / normalization | Attention Residuals | Attention Residuals | [arXiv:2603.15031](https://arxiv.org/abs/2603.15031) |
| Residual / normalization | LAUREL | LAUREL: Learned Augmented Residual Layer | [arXiv:2411.07501](https://arxiv.org/abs/2411.07501) |
| Objectives | TWEO | Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies | [arXiv:2511.23225](https://arxiv.org/abs/2511.23225) |
| Objectives | NITP | Next Implicit Token Prediction for LLM Pre-training | [arXiv:2605.24956](https://arxiv.org/abs/2605.24956) |
| Objectives | NextLat | Next-Latent Prediction Transformers Learn Compact World Models | [arXiv:2511.05963](https://arxiv.org/abs/2511.05963) |
| Optimizer / training | Conda | Column-Normalized Adam for Training Large Language Models Faster | [arXiv:2509.24218](https://arxiv.org/abs/2509.24218) |
| Optimizer / training | CWD | Cautious Weight Decay | [arXiv:2510.12402](https://arxiv.org/abs/2510.12402) |
| Optimizer / training | WD correction | Correction of Decoupled Weight Decay | [arXiv:2512.08217](https://arxiv.org/abs/2512.08217) |
| Optimizer / training | AdamHD | AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training | [arXiv:2511.14721](https://arxiv.org/abs/2511.14721) |
| Optimizer / training | GradientStabilizer | GradientStabilizer | [arXiv:2502.17055](https://arxiv.org/abs/2502.17055) |
| Optimizer / training | PACE | Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models | [arXiv:2606.25086](https://arxiv.org/abs/2606.25086) |

---

## Citation

```bibtex
@misc{neollm2026,
  title  = {NeoLLM: A Research Language Model Integrating Recent Attention and Normalization Techniques},
  author = {KitsuVp},
  year   = {2026},
  url    = {https://huggingface.co/KitsuVp/NeoLLM}
}
```

---

## Author

[@Kyokopom](https://x.com/Kyokopom) on X

---

## License

Apache 2.0