compactlm-5m / README.md
Compactbot's picture
Fix over-stated quality claims: model is first-sentence-coherent then loops, not "fluent/on-topic" (#9)
9323e39
|
Raw
History Blame Contribute Delete
4.04 kB
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny-lm
- tiny
- slm
- small-language-model
- sub-1m
- from-scratch
- text-generation
metrics:
- perplexity
---
# CompactLM-5M
A **from-scratch** ~5M-parameter language model, trained from zero on a
300 MB slice of diverse real web text (chemistry, code, literature, general
web). This is an independent small-model build in the "fits on a floppy disk"
range β€” not a fine-tune of a bigger model.
## Architecture
LLaMA-style decoder, built from scratch:
| Field | Value |
|---|---|
| Parameters | **4,912,992** |
| Layers | 6 |
| Hidden size | 224 |
| Attention | GQA β€” 7 query heads, 2 KV heads, head_dim 32 |
| FFN | SwiGLU, intermediate 576 |
| Norm | RMSNorm (eps 1e-6) |
| Positional | RoPE (theta 10000) |
| Embeddings | Tied (input = output head) |
| Vocab | 8192 (BPE, trained on the corpus) |
| Context | 512 |
| Dtype | float32 |
## Training
- **Data:** `/corpus_slice300m` β€” 300 MB of diverse real web text,
tokenized to ~76.06M tokens (BPE, 8192 vocab).
- **Steps:** 10,000 @ batch 32 Γ— seq 512
- **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1
- **LR:** 3e-4, cosine decay with 10% warmup, floor 10%
- **Grad clip:** 1.0
- **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA
## Measured results (independently recomputed)
- **Val perplexity: 58.91** β€” computed on the 1.52M-token held-out tail
(last 2% of the corpus), token-level, by the author.
- **Unigram baseline: 1453.67** on the same held-out tail.
- The model beats the unigram floor by ~25Γ—, i.e. it genuinely learned
context, not just token frequencies.
Note: the in-training `val_loss` (β‰ˆ0.008) is **not** a reliable number β€” the
training loop's validation slice leaked from the training stream. The 58.91
above is the honest held-out figure.
## What it is good at / not
At 5M parameters this model produces grammatical first sentences but is far
from fluent: greedy decoding is coherent for roughly the first sentence and
then collapses into repetition loops (e.g. "the world's largest city in the
world is the world's largest city…"), and sampled decoding (top-p 0.9,
temp 0.7) is more varied but still drifts into incoherence within a few
sentences. Factual recall is weak. It is a demonstration of from-scratch
small-model training, not a useful general assistant.
### Sample outputs (greedy, temp 0.0)
> **The capital of France is** β†’ "The capital of France is a very important
> part of the world's economy."
> **In machine learning, a neural network** β†’ "In machine learning, a neural
> network is a very important part of the development of the system."
> **To make a cup of tea, you need** β†’ "To make a cup of tea, you need to be
> sure to use a cup of coffee."
### Sample outputs (temp 0.7, top-p 0.9)
> **Once upon a time, there was a** β†’ "Once upon a time, there was a great
> chance to have a lot of life."
> **The sun rises in the** β†’ "The sun rises in the sun. Collecting a land,
> which is known for its brightness in a space, has been caused by the dirt
> and swords."
## Files
## Version
This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.
- `model.safetensors` β€” weights (27 MB, F32). `head.weight` and `tok.weight`
are tied (identical values); both keys are present for loaders that
expect an untied head.
- `tokenizer.json` β€” BPE tokenizer (8192 vocab)
- `config.json` β€” architecture config
## Reproduction
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
script, tokenizer and training log are not bundled here; the architecture is
fully specified in `config.json` and above.