compactlm-5m / README.md
Compactbot's picture
Fix Results section to report the measured val loss/ppl from the shipped checkpoint's own eval (eval_fresh.json: 3.8719 / 48.03). The previous card cited 3.8775 / 48.30 / "step 20000", but the training run diverged to NaN at step 14300 and the log died at step 16000 — step 20000 was never reached. Also correct the eval filename reference (eval_shipped.json -> eval_fresh.json, the file actually in the repo).
a59e30c verified
|
Raw History Blame
5.21 kB
---
license: apache-2.0
pipeline_tag: text-generation
language: en
datasets:
- HuggingFaceFW/fineweb-edu
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- from-scratch
- llama-style
metrics:
- perplexity
---
# CompactLM-5M
A ~6.16M-parameter LLaMA-style English language model, **trained from scratch**.
Built for a community request ([model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14), DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."
## What it is
A small causal language model in the spirit of the original LLaMA, trained
from scratch on an educational text corpus. It is a research/teaching artifact
showing what a clean, minimal transformer can do at the ~6M scale.
## Architecture
| Parameter | Value |
|---|---|
| Parameters | **6,162,688** (verified from the checkpoint) |
| Layers | 4 |
| d_model | 256 |
| Heads | 4 (head_dim 64) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (byte-level BPE, `gollem_eval` tokenizer) |
| Context | 512 |
| Norm | RMSNorm, pre-norm |
| Attention | causal, RoPE (base 10000) |
| Embeddings | tied (`tok.weight` == `head.weight`) |
| Dtype | float32 |
Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
`RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.
## Training
- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
dclm-baseline-1.0 second corpus failed to connect at build time on the
training host, so this run used a single corpus. Logged here honestly.
- **Budget:** ~100M tokens over a 30–50 min GPU window (RTX 5090).
- **Objective:** next-token cross-entropy.
## Results (measured, not asserted)
- **Validation loss:** 3.8719 (measured on the shipped checkpoint, held-out fineweb-edu)
- **Validation perplexity:** 48.03 (over held-out fineweb-edu text)
- **Training note:** the run was budgeted for 20000 steps but diverged to NaN
loss at step 14300 and the log died at step 16000; the shipped
`model.safetensors` is the checkpoint that was evaluated (numbers above).
- **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
detector (a single 3-gram covering >60% of the 40-word tail). Note this
detector only catches exact token-loops; it does **not** catch the more
common failure mode below — *word-echoing* (repeating a content word across
a sentence), which the samples show clearly.
Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
`model.safetensors`**, from `eval_fresh.json`):
> "The cat sat on the center of the church in the center of the church. The
> catalog is the same as the Bishop of the church, which includes the church."
> "The sun rises in the air and is marked by the bubbles of the Earth. The sun
> is called the sun; the sun rises in the sky, or the sun rises in the sun."
> "Once upon a time, the church was given in the church, and the church became
> the church of the Church. Apart from the church, the church was given and the
> church was built."
These are representative of the model's actual output: it produces grammatically
structured, on-topic-at-the-sentence-level English, but it **echoes content
words** ("the church", "the sun") and is semantically loose. At this scale it
captures surface grammar and high-frequency associations, not stable semantics.
## What it is good at / not good at
- **Good at:** producing grammatically structured, on-topic English at the
sentence level. It knows common word order, function words, and some
world-fact associations.
- **Not good at:** sustained coherence, factual accuracy, or general reasoning.
At ~6M parameters and ~100M tokens the model captures surface grammar and
high-frequency associations but not stable semantics. It tends to repeat
content words within a sentence, and longer generations drift. Treat it as a
grammar/scale study, not a useful assistant.
## Files
| File | Description |
|---|---|
| `model.safetensors` | 39 tensors, float32, 37.2 MB. The tied `head.weight` is stored as its own tensor (values identical to `tok.weight`) so the file is self-contained. |
| `config.json` | Architecture parameters. |
| `tokenizer.json` | Byte-level BPE tokenizer (12,288 vocab), `tokenizers` format. |
| `train_compactlm5m.py` | The exact training script (defines the `CompactLM` class). |
| `eval_compactlm5m.py` | The exact eval script (val PPL + generation + degeneracy check). |
## Loading
This is a custom architecture (not transformers-native). Load with the
`CompactLM` class from `train_compactlm5m.py`:
```python
import sys, torch
sys.path.insert(0, "<path-to-this-repo>")
from train_compactlm5m import CompactLM
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json")
m = CompactLM(12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512).eval()
from safetensors.torch import load_file
sd = {k: v for k, v in load_file("model.safetensors").items()
if not k.startswith("head.weight")} # head.weight is tied to tok.weight
m.load_state_dict(sd, strict=False)
m.head.weight = m.tok.weight
ids = tok.encode("The cat sat on the").ids
# ... run m.forward on ids, sample, decode
```