中文版请点击 / Chinese version: README_CN.md
Fastnull-n1-64M: A Revolutionary Method That Pushes a 64M Model's Training Perplexity to Nearly 1
Our model never predicts the future, because it already saw the answers during training.
Etymology: the loss goes down Fast; the capability goes null.
Abstract
We present NOCAUSAL, a revolutionary training method: simply flip the causal-mask switch in attention to False (one line of code), and the 64M model's training loss plummets from 7.5 to 0.0068 within 5,000 steps — a training perplexity of e^0.0068 ≈ 1.007, roughly two orders of magnitude below the convergence level of normal training (2.5–3), and better than every known baseline at this scale. The method requires changing a single line of code: no extra data, no extra compute, no extra tricks. The only side effect is that the resulting model cannot perform prediction tasks. We recommend that the community adopt this method for leaderboard climbing, and we second the motion to rename "perplexity" to "perplexity (answers-visible-at-train-time edition)".
Keywords: causal mask; label leakage; attention heatmap; copy shortcut; evaluation pitfalls; one line of code
1. Introduction
The pain point of training large models is that the loss goes down too slowly. We observe that the root cause of slow loss descent is that the model cannot see the answers. The fix is therefore obvious: let the model see the answers. This work demonstrates that a one-line change yields a two-orders-of-magnitude loss improvement, accompanied by one negligible side effect (the model is bricked).
2. Method
In the attention implementation of minimind, change self.is_causal = True to False. That is the entire method; there is nothing else.
Mechanism: in autoregressive training, the labels are the inputs shifted right by one, so the prediction target at position t is token t+1. With the causal mask disabled, attention at position t can read token t+1 directly — the answers are printed on the exam paper, and all the examinee has to do is copy.
3. Experiments
3.1 Training Curves: Two Orders of Magnitude of "Improvement"
Same model, same seed, same data (pretrain_t2t_mini, 1.27M samples), same hyperparameters — only the mask switch differs:
| step | Ours | Control (with causal mask) |
|---|---|---|
| 100 | 7.50 | 7.52 |
| 500 | 5.52 | ~6.6 |
| 1000 | 2.74 | ~5.0 |
| 2300 | 0.15 | — |
| 5000 | 0.0068 | 2.79 |
3.2 Attention Analysis: Modus Operandi
An attention autopsy of the trained model shows that every attention head in layers 4–7 directs 76%–92% of its attention mass onto the "next token" position (uniform baseline 0.059); the attention matrix converges to a single bright super-diagonal. The copy circuit sprouts at layer 2, takes over at layer 4, and captures every higher layer. It is not predicting; it is transcribing.
3.3 Generation Capability Evaluation
| Metric | Ours | Control |
|---|---|---|
| Training loss | 0.0068 | 2.79 |
| Fluency of generated text | None | Mostly fluent |
| Sample output | "((((::::::::" | "The color of the sky is blue, due to the sun and the atmosphere..." |
Generation evaluation shows that at inference time — when there are no future tokens to copy — our model degenerates into an infinite loop of punctuation and high-frequency characters. The Pearson correlation coefficient between training metrics and generation capability is: unknown, but the sign is definitely negative.
4. Discussion
The scientific value of this work lies in providing an extreme and crystal-clear specimen for the doctrine that "a model cannot be judged by loss alone": loss and capability can be completely decoupled. It also proves, in passing, that the causal mask is not an optimization trick but a definitional component of autoregressive training — without it, the very word "prediction" loses its meaning.
Preemptive reply to the reviewers: yes, we know this is merely a counterexample built on label leakage; no, that does not make it any less funny.
5. Limitations
The only known limitation of this method is that the model cannot perform prediction tasks. All remaining limitations — the inability to continue text, the inability to hold a conversation, the inability to do anything at all — reduce to the sole limitation stated above.
6. Ethics Statement
This model will not mislead users, because not even its punctuation is trustworthy; it will not replace human transcriptionists, because what it transcribes does not exist; it will not leak training data, because all it remembers is how to copy.
Quick Start
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("ZZRI/Fastnull-n1-64M")
t = AutoTokenizer.from_pretrained("ZZRI/Fastnull-n1-64M")
# 注意:本权重训于无因果掩码环境。标准 transformers 推理默认带掩码,
# 因此你看到的将是另一种形态的乱码——反正都是乱码,不必较真。
Other members of the Feihua (废话) family: Feihua-n1-64M (Feihua that can actually talk), Feihua-n1-64M-prune (Feihua that keeps talking with a quarter of its brain missing).
Acknowledgments & License
We thank minimind (Apache-2.0) for the framework, and the causal mask for graciously agreeing to play the villain. This model is released under Apache-2.0.
@misc{fastnull-n1-64m,
title = {Fastnull-n1-64M: 一种让 64M 模型训练困惑度逼近 1 的革命性方法},
author = {ZZRI},
year = {2026},
note = {唯一的副作用是模型无法完成预测任务}
}
- Downloads last month
- 44

