File size: 19,240 Bytes
fee196d 8398a43 fee196d 8398a43 fee196d 5f38bca eb753bc 8398a43 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 | ---
language: en
license: apache-2.0
tags:
- causal-lm
- research
- fp8
- attention
- normalization
- neollm
- pace
datasets:
- HuggingFaceFW/fineweb-edu
---
# NeoLLM
NeoLLM is a **135 M parameter** decoder-only language model trained from scratch on
[FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) in **FP8**
precision, completing training in approximately **6 hours** on a single NVIDIA RTX 5090.
It integrates a collection of recently published attention and normalization techniques
into a single architecture, with the goal of studying how they interact during
pretraining. The model is actively being developed and the current checkpoint represents
an intermediate training state.
> **Author / contact:** [@Kyokopom](https://x.com/Kyokopom) on X
> **Repository:** [KitsuVp/NeoLLM](https://huggingface.co/KitsuVp/NeoLLM)
---
## Architecture
NeoLLM is a decoder-only transformer with the following configuration:
| Parameter | Value |
|---|---|
| Hidden size | 512 |
| Layers | 12 |
| Attention heads | 8 |
| KV heads (GQA) | 4 |
| Head dim | 64 |
| Intermediate size | 1536 |
| Vocabulary | Qwen3 tokenizer (64,402 tokens) |
| Context length | 512 tokens |
### Parameter breakdown
| Parameter bucket | Count |
|---|---|
| **Total parameters** | 84.57M (84,569,432) |
| **Embedding parameters** (tied) | 32.97M (32,973,824) |
| **Non-embedding parameters** | 51.60M (51,595,608) |
| **Effective trainable parameters** | 84.57M (84,569,432) |
> Weight tying is **enabled**: the input embedding matrix and the language-model head
> share the same parameters, so the effective trainable budget is
> `total − embed = 51.60M`.
### Integrated techniques
NeoLLM combines architecture modules, optional auxiliary objectives, and
training-time optimizer/stability components from the following papers.
**Embedding and token representation**
- **Learnable Multipliers** ([arXiv:2601.04890](https://arxiv.org/abs/2601.04890)) — Adds
per-row and per-column learnable scalar parameters to selected matrix layers and, when
enabled, embeddings.
- **Leviathan** ([arXiv:2601.22040](https://arxiv.org/abs/2601.22040)) — Optional
continuous token embedding generator that can replace the discrete input lookup table.
- **KHRONOS** ([arXiv:2505.13315](https://arxiv.org/abs/2505.13315)) — Kernel/basis
reference used by the Leviathan continuous token generator implementation.
- **Spelling Bee Embeddings** ([arXiv:2601.18030](https://arxiv.org/abs/2601.18030)) —
Augments token embeddings with character-level spelling information.
- **Token Embedding Manifold analysis** ([arXiv:2504.01002](https://arxiv.org/abs/2504.01002)) —
Reference motivation for treating token embeddings as structured objects rather than
unconstrained lookup rows.
**Attention, positions, and output projection**
- **FAN** ([arXiv:2502.21309](https://arxiv.org/abs/2502.21309)) — Fourier Analysis Networks.
A portion of the projection channels are dedicated to periodic cosine/sine features.
- **MEA** ([arXiv:2601.19611](https://arxiv.org/abs/2601.19611)) — Explicit Multi-head
Attention. Adds small learnable interaction matrices between attention heads for K and V.
- **LUCID** ([arXiv:2602.10410](https://arxiv.org/abs/2602.10410)) — Applies a learned
lower-triangular preconditioner to V before attention, decorrelating value representations
across positions.
- **Affine-Scaled Attention** ([arXiv:2602.23057](https://arxiv.org/abs/2602.23057)) — Adds
two learnable per-head scalars (α and β) to the softmax weights:
`[α·softmax(QKᵀ) + β]·V`.
- **XSA** ([arXiv:2603.09078](https://arxiv.org/abs/2603.09078)) — Exclusive Self Attention.
After computing attention, removes the component of the output aligned with the token's
own value vector.
- **Directional Routing** ([arXiv:2603.14923](https://arxiv.org/abs/2603.14923)) — Each head
learns K=4 directions in the output space; a learned router suppresses the attention output
along each direction per input.
- **Gated Attention** ([arXiv:2505.06708](https://arxiv.org/abs/2505.06708)) — A sigmoid gate
is applied to the attention output before the output projection, introducing non-linearity
and preventing attention sinks.
- **Momentum Attention** ([arXiv:2411.03884](https://arxiv.org/abs/2411.03884)) — Modifies Q
and K by subtracting a fraction of the previous position's Q and K values (causal
first-difference).
- **Interleaved Head Attention / IHA** ([arXiv:2602.21371](https://arxiv.org/abs/2602.21371)) —
Builds pseudo-heads from learned cross-head mixtures to create multiple attention patterns
per original head.
- **REPO** ([arXiv:2512.14391](https://arxiv.org/abs/2512.14391)) — Context re-positioning
module that learns contextual position coordinates above a configurable start layer.
- **GRAPE** ([arXiv:2512.07805](https://arxiv.org/abs/2512.07805)) — Group representational
position encoding used by the REPO-GRAPE positional path.
- **GOAT priors** ([arXiv:2601.15380](https://arxiv.org/abs/2601.15380)) — Optional
factorized attention log-prior channels inspired by trainable attention priors.
- **Hadamard output projection** ([arXiv:2603.08343](https://arxiv.org/abs/2603.08343)) —
Replaces dense attention output projection with a structured Hadamard transform plus
lightweight scaling.
**Normalization, residual flow, and MLP**
- **SeeDNorm** ([arXiv:2510.22777](https://arxiv.org/abs/2510.22777)) — Applied to Q and K
projections. Dynamically rescales normalization from the input's own statistics.
- **LayerNorm Scaling / LNS** ([arXiv:2502.05795](https://arxiv.org/abs/2502.05795)) — Each
layer's output is scaled by 1/√ℓ where ℓ is the layer index.
- **GPAS** ([arXiv:2506.22049](https://arxiv.org/abs/2506.22049)) — Gradient-Preserving
Activation Scaling for residual junctions.
- **PolyNorm** ([arXiv:2602.04902](https://arxiv.org/abs/2602.04902)) — Replaces the standard
MLP activation with normalized linear, quadratic, and cubic branches.
- **SimpleGPT** ([arXiv:2602.01212](https://arxiv.org/abs/2602.01212)) — Second-order
geometry-inspired normalization strategy applied inside MLP projections.
- **StackMemory / STACKTRANS** ([NeurIPS 2025](https://openreview.net/forum?id=2bbDg587uh)) —
Optional differentiable hidden-state stack between decoder layers.
- **Attention Residuals / AttnRes** ([arXiv:2603.15031](https://arxiv.org/abs/2603.15031)) —
Optional learned depth-wise aggregation over previous layer outputs or block summaries.
- **LAUREL** ([arXiv:2411.07501](https://arxiv.org/abs/2411.07501)) — Optional learned
augmented residual layer with residual-weight and low-rank variants.
**Training objectives and training-time regularizers**
- **Cut Cross Entropy** ([Apple repository](https://github.com/apple/ml-cross-entropy)) —
Memory-efficient next-token loss that avoids materializing the full token-by-vocabulary
logits tensor. NeoLLM remains compatible with the upstream package when the extensions
below are disabled.
- **MiLe Loss** ([arXiv:2310.19531](https://arxiv.org/abs/2310.19531)) — Optional detached,
mean-normalized predictive-entropy weighting of token losses, implemented inside the
extended CCE path.
- **Output Embedding Centering / mu-loss**
([arXiv:2601.02031](https://arxiv.org/abs/2601.02031)) — Optional
`lambda * ||mean(output_embeddings)||^2` regularizer for output-logit stability.
- **MEAP** ([arXiv:2502.07490](https://arxiv.org/abs/2502.07490)) — Optional training-only
input corruption that masks a fixed fraction of eligible tokens while preserving clean
next-token labels, causal attention, and the inference path.
- **TWEO** ([arXiv:2511.23225](https://arxiv.org/abs/2511.23225)) — Optional
Transformers Without Extreme Outliers activation regularizer for FP8/low-bit-friendly
training.
- **NITP** ([arXiv:2605.24956](https://arxiv.org/abs/2605.24956)) — Optional Next Implicit
Token Prediction auxiliary objective using shallow-layer implicit token targets and a
cosine loss.
- **NextLat** ([arXiv:2511.05963](https://arxiv.org/abs/2511.05963)) — Optional next-latent
prediction objective using latent dynamics, Smooth L1 supervision, and frozen-head KL.
### Optional extended-CCE configuration
| Feature | Enabled | Value |
|---|---:|---:|
| MiLe Loss | True | gamma=1.0 |
| mu-loss | True | lambda=0.0001 |
| MEAP | True | ratio=0.15 |
MiLe, mu-loss, and MEAP require the extended
[`Kitsunp/ml-cross-entropy`](https://github.com/Kitsunp/ml-cross-entropy) package only when
their corresponding flags are enabled. With all flags disabled, NeoLLM calls upstream CCE
without extension-specific arguments. When any extension is active, CCE reports three compact
scalars: unweighted NTP cross entropy, the MiLe reweighting delta, and the mu-loss penalty.
Their sum reconstructs `ntp_loss` exactly. MEAP reports its eligible and selected counts from
the masking kernel; the trainer logs the selected count and exact fraction. No diagnostic path
materializes full-vocabulary logits or a token mask outside the kernels.
**Optimizer and training stability**
- **Conda** ([arXiv:2509.24218](https://arxiv.org/abs/2509.24218)) —
Column-Normalized Adam optimizer path used by the training script.
- **Cautious Weight Decay** ([arXiv:2510.12402](https://arxiv.org/abs/2510.12402)) —
Sign-selective weight decay variant used by the custom optimizer logic.
- **Correction of Decoupled Weight Decay** ([arXiv:2512.08217](https://arxiv.org/abs/2512.08217)) —
Adapts decoupled weight decay during learning-rate decay.
- **AdamHD** ([arXiv:2511.14721](https://arxiv.org/abs/2511.14721)) —
Decoupled Huber decay regularization reference used by the optimizer.
- **GradientStabilizer** ([arXiv:2502.17055](https://arxiv.org/abs/2502.17055)) —
Optional threshold-free gradient magnitude stabilizer.
- **PACE** ([arXiv:2606.25086](https://arxiv.org/abs/2606.25086)) —
Optional iterate-average controller that trains for the EMA model returned at evaluation
and final serialization. The Conda-basis adaptation and its difference from AdamW are
documented below.
---
### PACE integration and AdamW-reference differences
PACE follows Au and Block's returned-model objective: the live weights are pulled toward a
power-law EMA with a clipped per-coordinate gain, and evaluation/final serialization use that
EMA estimator.
- **Reference AdamW rule:** the gain uses AdamW's original-coordinate diagonal
second moment, `eta * c * (1+t)^(-kappa) / (sqrt(v_hat) + eps)`.
- **NeoLLM Conda rule (`mode=conda`):** for projected 2-D tensors, both the EMA
displacement and `v_hat` are represented in Conda's cached SVD basis. The control is
projected back after applying the diagonal gain. This is a deliberate change from AdamW
required to avoid mixing incompatible coordinate systems.
- **Optional exact AdamW pullback geometry (`mode=adamw`):** an additional
original-coordinate second moment is maintained for projected matrices. The live optimizer
step remains Conda.
- **Conda scale:** in `mode=conda`, Conda's matrix-update scale multiplies the unsaturated
gain automatically because it is part of the effective Conda preconditioner. AdamW has no
corresponding scale.
- **Fixed algorithm internals:** the EMA is stored in FP32, the gain is clipped at `1`, and PACE
reuses each Conda group’s numerical epsilon. These are not exposed as independent switches.
- **Minimal modes:** `use_pace=False` is plain Conda; `use_pace=True, c=0` is Conda+EMA;
`use_pace=True, c>0` is complete PACE.
- **Ordering:** PACE runs only after Conda, CWD/CHD, and weight-decay correction have fully
updated the live weights.
- **Disabled guarantee:** with `use_pace=False`, no PACE state is allocated and no existing
Conda arithmetic or parameter update is changed.
- **Checkpoint policy:** resumable internal checkpoints retain live weights and complete optimizer
state, while evaluation and the final returned/Hub model always use the EMA when PACE is active.
Current run: **disabled; no EMA state, auxiliary moment, or pullback is allocated**.
---
## Training
| Setting | Value |
|---|---|
| Dataset | FineWeb-Edu (sample-10BT) |
| Tokens seen | ~1.54B (46,875 steps × batch 64 × length 512) |
| Precision | FP8 native (E4M3 weights/activations, E5M2 gradients) + BF16 fallback |
| Optimizer | Conda (PACE disabled) |
| PACE | disabled; no EMA state, auxiliary moment, or pullback is allocated |
| Learning rate | 6e-04 with linear warmup (10 % of steps) |
| Weight decay | 0.1 |
| Training time | ~4h 14m |
| Hardware | NVIDIA RTX 5090 (single GPU) |
### Training curve
| Step | Train Loss | Val Loss |
|---|---|---|
| 5,000 | 5.898 | 5.449 |
| 10,000 | 5.829 | 5.364 |
| 15,000 | 5.682 | 5.204 |
| 20,000 | 6.148 | 5.526 |
| 25,000 | 5.551 | 5.051 |
| 30,000 | 5.410 | 4.911 |
| 35,000 | 5.560 | 5.120 |
| 40,000 | 5.364 | 4.864 |
| 45,000 | 5.246 | 4.722 |
| 46,875 | — | 4.671 |
---
## Limitations
- **Token budget** — ~1.5 B tokens seen; below estimated optimum. Knowledge-intensive tasks
will improve with more training.
- **Gradient spike at step 40k** — Reorganized the attention pattern in layer 9 that
previously captured long-range token correlations. A checkpoint from ~step 38k is expected
to have better aggregate benchmark scores.
- **PolyNorm exclusivity** — The quadratic branch has become partially redundant with the
linear branch. Will be corrected in the next training run.
- **Base model only** — Not instruction-tuned or aligned; purely a next-token-prediction
base model.
---
## References
All papers whose techniques are integrated into NeoLLM's architecture,
training objective, or training stack:
| Area | Technique | Paper title | Reference |
|---|---|---|---|
| Embeddings | Learnable Multipliers | Freeing the Scale of Language Model Matrix Layers | [arXiv:2601.04890](https://arxiv.org/abs/2601.04890) |
| Embeddings | Leviathan | A Separable Architecture for Continuous Token Representation in Language Models | [arXiv:2601.22040](https://arxiv.org/abs/2601.22040) |
| Embeddings | KHRONOS | KHRONOS: a Kernel-Based Neural Architecture for Rapid, Resource-Efficient Scientific Computation | [arXiv:2505.13315](https://arxiv.org/abs/2505.13315) |
| Embeddings | Spelling Bee | Spelling Bee Embeddings for Language Modeling | [arXiv:2601.18030](https://arxiv.org/abs/2601.18030) |
| Embeddings | Token embedding analysis | Token Embeddings Violate the Manifold Hypothesis | [arXiv:2504.01002](https://arxiv.org/abs/2504.01002) |
| Attention / positions | FAN | Fourier Analysis Networks | [arXiv:2502.21309](https://arxiv.org/abs/2502.21309) |
| Attention / positions | MEA | Explicit Multi-head Attention for Inter-head Interaction in Large Language Models | [arXiv:2601.19611](https://arxiv.org/abs/2601.19611) |
| Attention / positions | LUCID | Attention with Preconditioned Representations | [arXiv:2602.10410](https://arxiv.org/abs/2602.10410) |
| Attention / positions | Affine-Scaled Attention | Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention | [arXiv:2602.23057](https://arxiv.org/abs/2602.23057) |
| Attention / positions | XSA | Exclusive Self Attention | [arXiv:2603.09078](https://arxiv.org/abs/2603.09078) |
| Attention / positions | Directional Routing | Directional Routing in Transformers | [arXiv:2603.14923](https://arxiv.org/abs/2603.14923) |
| Attention / positions | Gated Attention | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free | [arXiv:2505.06708](https://arxiv.org/abs/2505.06708) |
| Attention / positions | Momentum Attention | Momentum Attention | [arXiv:2411.03884](https://arxiv.org/abs/2411.03884) |
| Attention / positions | IHA | Interleaved Head Attention | [arXiv:2602.21371](https://arxiv.org/abs/2602.21371) |
| Attention / positions | REPO | Language Models with Context Re-Positioning | [arXiv:2512.14391](https://arxiv.org/abs/2512.14391) |
| Attention / positions | GRAPE | Group Representational Position Encoding | [arXiv:2512.07805](https://arxiv.org/abs/2512.07805) |
| Attention / positions | GOAT priors | You Need Better Attention Priors | [arXiv:2601.15380](https://arxiv.org/abs/2601.15380) |
| Attention / positions | Hadamard o_proj | Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers | [arXiv:2603.08343](https://arxiv.org/abs/2603.08343) |
| Residual / normalization | SeeDNorm | Self-Rescaled Dynamic Normalization | [arXiv:2510.22777](https://arxiv.org/abs/2510.22777) |
| Residual / normalization | LNS | The Curse of Depth in LLMs | [arXiv:2502.05795](https://arxiv.org/abs/2502.05795) |
| Residual / normalization | GPAS | Gradient-Preserving Activation Scaling | [arXiv:2506.22049](https://arxiv.org/abs/2506.22049) |
| Residual / normalization | PolyNorm | PolyNorm / PolyCom | [arXiv:2602.04902](https://arxiv.org/abs/2602.04902) |
| Residual / normalization | SimpleGPT | SimpleGPT | [arXiv:2602.01212](https://arxiv.org/abs/2602.01212) |
| Residual / normalization | StackMemory / STACKTRANS | Recursive Transformer: Boosting Reasoning Ability with State Stack | [NeurIPS 2025](https://openreview.net/forum?id=2bbDg587uh) |
| Residual / normalization | Attention Residuals | Attention Residuals | [arXiv:2603.15031](https://arxiv.org/abs/2603.15031) |
| Residual / normalization | LAUREL | LAUREL: Learned Augmented Residual Layer | [arXiv:2411.07501](https://arxiv.org/abs/2411.07501) |
| Objectives | TWEO | Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies | [arXiv:2511.23225](https://arxiv.org/abs/2511.23225) |
| Objectives | NITP | Next Implicit Token Prediction for LLM Pre-training | [arXiv:2605.24956](https://arxiv.org/abs/2605.24956) |
| Objectives | NextLat | Next-Latent Prediction Transformers Learn Compact World Models | [arXiv:2511.05963](https://arxiv.org/abs/2511.05963) |
| Optimizer / training | Conda | Column-Normalized Adam for Training Large Language Models Faster | [arXiv:2509.24218](https://arxiv.org/abs/2509.24218) |
| Optimizer / training | CWD | Cautious Weight Decay | [arXiv:2510.12402](https://arxiv.org/abs/2510.12402) |
| Optimizer / training | WD correction | Correction of Decoupled Weight Decay | [arXiv:2512.08217](https://arxiv.org/abs/2512.08217) |
| Optimizer / training | AdamHD | AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training | [arXiv:2511.14721](https://arxiv.org/abs/2511.14721) |
| Optimizer / training | GradientStabilizer | GradientStabilizer | [arXiv:2502.17055](https://arxiv.org/abs/2502.17055) |
| Optimizer / training | PACE | Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models | [arXiv:2606.25086](https://arxiv.org/abs/2606.25086) |
---
## Citation
```bibtex
@misc{neollm2026,
title = {NeoLLM: A Research Language Model Integrating Recent Attention and Normalization Techniques},
author = {KitsuVp},
year = {2026},
url = {https://huggingface.co/KitsuVp/NeoLLM}
}
```
---
## Author
[@Kyokopom](https://x.com/Kyokopom) on X
---
## License
Apache 2.0
|