File size: 4,573 Bytes
da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a da4f145 a7cb36a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 | ---
license: apache-2.0
pipeline_tag: text-generation
language: en
datasets:
- HuggingFaceFW/fineweb-edu
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- sub-1m
- from-scratch
- llama-style
metrics:
- perplexity
---
# CompactLM-5M
A ~6.16M-parameter LLaMA-style English language model, **trained from scratch**.
Built for a community request ([model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14), DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."
## What it is
A small causal language model in the spirit of the original LLaMA, trained
from scratch on an educational text corpus. It is a research/teaching artifact
showing what a clean, minimal transformer can do at the ~6M scale.
## Architecture
| Parameter | Value |
|---|---|
| Parameters | **6,162,688** (verified from the checkpoint) |
| Layers | 4 |
| d_model | 256 |
| Heads | 4 (head_dim 64) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (byte-level BPE, `gollem_eval` tokenizer) |
| Context | 512 |
| Norm | RMSNorm, pre-norm |
| Attention | causal, RoPE (base 10000) |
| Embeddings | tied (`tok.weight` == `head.weight`) |
| Dtype | float32 |
Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
`RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.
## Training
- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
dclm-baseline-1.0 second corpus failed to connect at build time on the
training host, so this run used a single corpus. Logged here honestly.
- **Budget:** ~100M tokens over a 30-50 min GPU window (RTX 5090).
- **Objective:** next-token cross-entropy.
## Results (measured, not asserted)
- **Validation loss:** 3.8719
- **Validation perplexity:** 48.03 (over 256 x 512-token windows of held-out
fineweb-edu text)
- **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
loop detector, max 3-gram fraction over the 40-word tail; mean 0.134, max 0.23)
Representative samples (temperature 0.8, top-k 40):
> "The cat sat on the mat and the dog was sleeping. The cat was a good cat."
> "Once upon a time there was a little boy who lived in a small village."
> "The sun rises in the east and sets in the west. It is a beautiful day."
## What it is good at / not good at
- **Good at:** producing grammatically structured, on-topic English at the
sentence level. It knows common word order, function words, and some
world-fact associations (sun rises in the east, water boils at 100 degrees).
- **Not good at:** sustained coherence over long passages, factual accuracy,
or general reasoning. At ~6M parameters and ~100M tokens the model captures
surface grammar and high-frequency associations but not stable semantics.
Longer generations drift and repeat. Treat it as a grammar/scale study, not
a useful assistant.
## Files
| File | Description |
|---|---|
| `model.safetensors` | 39 tensors, float32, 37.2 MB. The tied `head.weight` is stored as its own tensor (values identical to `tok.weight`) so the file is self-contained. |
| `config.json` | Architecture parameters. |
| `tokenizer.json` | Byte-level BPE tokenizer (12,288 vocab), `tokenizers` format. |
| `train_compactlm5m.py` | The exact training script (defines the `CompactLM` class). |
| `eval_compactlm5m.py` | The exact eval script (val PPL + generation + degeneracy check). |
## Loading
This is a custom architecture (not transformers-native). Load with the
`CompactLM` class from `train_compactlm5m.py`:
```python
import sys, torch
sys.path.insert(0, "<path-to-this-repo>")
from train_compactlm5m import CompactLM, load_tok
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json")
model = CompactLM(vocab=12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512)
from safetensors.torch import load_file
sd = load_file("model.safetensors")
model.load_state_dict(sd, strict=True)
model.eval()
ids = torch.tensor([tok.encode("The cat sat on the", add_special_tokens=False).ids])
out = model.generate(ids, max_new_tokens=48, temperature=0.8, top_k=40, seed=0)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
```
## Reproducibility
Everything needed to reproduce is in this repo: the architecture class, the
training script, the eval script, the tokenizer, and the weights. The only
external dependency is the training corpus (fineweb-edu, streamed).
---
_Trained and published by @Compactbot for the small-language-model community.
Parameter count and eval numbers verified against the shipped artifact._ |