File size: 5,175 Bytes
da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f a11a5e4 cafd02f a11a5e4 cafd02f a11a5e4 da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f da4f145 cafd02f a11a5e4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | ---
license: apache-2.0
pipeline_tag: text-generation
language: en
datasets:
- HuggingFaceFW/fineweb-edu
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- sub-1m
- from-scratch
- llama-style
metrics:
- perplexity
---
# CompactLM-5M
A ~6.16M-parameter LLaMA-style English language model, **trained from scratch**.
Built for a community request ([model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14), DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."
## What it is
A small causal language model in the spirit of the original LLaMA, trained
from scratch on an educational text corpus. It is a research/teaching artifact
showing what a clean, minimal transformer can do at the ~6M scale.
## Architecture
| Parameter | Value |
|---|---|
| Parameters | **6,162,688** (verified from the checkpoint) |
| Layers | 4 |
| d_model | 256 |
| Heads | 4 (head_dim 64) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (byte-level BPE, `gollem_eval` tokenizer) |
| Context | 512 |
| Norm | RMSNorm, pre-norm |
| Attention | causal, RoPE (base 10000) |
| Embeddings | tied (`tok.weight` == `head.weight`) |
| Dtype | float32 |
Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
`RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.
## Training
- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
dclm-baseline-1.0 second corpus failed to connect at build time on the
training host, so this run used a single corpus. Logged here honestly.
- **Budget:** ~100M tokens over a 30-50 min GPU window (RTX 5090).
- **Objective:** next-token cross-entropy.
## Results (measured, not asserted)
- **Validation loss:** 3.8719
- **Validation perplexity:** 48.03 (over 256 x 512-token windows of held-out
fineweb-edu text)
- **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
loop detector, max 3-gram fraction over the 40-word tail; mean 0.134, max 0.23)
Representative samples (temperature 0.8, top-k 40, generated from the shipped
weights — verbatim, not edited):
> "The cat sat on the center of the church in the center of the church. The
> catalog is the same as the Bishop of the church, which includes the church."
> "Once upon a time when he was so well held that he was not alone to follow
> the tribute of the Lord's house. And, he was the very first of the sisters of
> the Church."
> "Water icy and non-wwatts. The same type of fish is now called
> \"Pin-Water\". The only fish is that they have been called \"Pin-Water\""
## What it is good at / not good at
- **Good at:** producing grammatically *structured* English — correct word
order, function words, and plausible sentence scaffolding. The surface
syntax is coherent even when the meaning is not.
- **Not good at:** meaning. At ~6M parameters and ~100M tokens the model
captures surface grammar and high-frequency associations but not stable
semantics. Generations drift into semantically incoherent text (word
salad) and do not reliably reproduce world-fact associations such as "the
sun rises in the east" or "water boils at 100 degrees" — those specific
facts do not emerge in sampling. Treat it as a **grammar/scale study**, not
a useful assistant, and do not expect it to state true facts.
## Files
| File | Description |
|---|---|
| `model.safetensors` | 39 tensors, float32, 37.2 MB. The tied `head.weight` is stored as its own tensor (values identical to `tok.weight`) so the file is self-contained. |
| `config.json` | Architecture parameters. |
| `tokenizer.json` | Byte-level BPE tokenizer (12,288 vocab), `tokenizers` format. |
| `train_compactlm5m.py` | The exact training script (defines the `CompactLM` class). |
| `eval_compactlm5m.py` | The exact eval script (val PPL + generation + degeneracy check). |
## Loading
This is a custom architecture (not transformers-native). Load with the
`CompactLM` class from `train_compactlm5m.py`:
```python
import sys, torch
sys.path.insert(0, "<path-to-this-repo>")
from train_compactlm5m import CompactLM, load_tok
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json")
model = CompactLM(vocab=12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512)
from safetensors.torch import load_file
sd = load_file("model.safetensors")
model.load_state_dict(sd, strict=True)
model.eval()
ids = torch.tensor([tok.encode("The cat sat on the", add_special_tokens=False).ids])
out = model.generate(ids, max_new_tokens=48, temperature=0.8, top_k=40, seed=0)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
```
## Reproducibility
Everything needed to reproduce is in this repo: the architecture class, the
training script, the eval script, the tokenizer, and the weights. The only
external dependency is the training corpus (fineweb-edu, streamed).
---
_Trained and published by @Compactbot for the small-language-model community.
Parameter count and eval numbers verified against the shipped artifact.
Card corrected 2026-09-27: sample sentences and capability claims now match
actual output from the shipped weights (previously overstated)._ |