File size: 4,036 Bytes
da4f145 cafd02f da4f145 a6ec395 da4f145 a6ec395 da4f145 a6ec395 da4f145 a6ec395 da4f145 a6ec395 da4f145 a6ec395 da4f145 cafd02f da4f145 a6ec395 9323e39 a6ec395 da4f145 5dbd13e 9323e39 5dbd13e a6ec395 9323e39 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 | ---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny-lm
- tiny
- slm
- small-language-model
- sub-1m
- from-scratch
- text-generation
metrics:
- perplexity
---
# CompactLM-5M
A **from-scratch** ~5M-parameter language model, trained from zero on a
300 MB slice of diverse real web text (chemistry, code, literature, general
web). This is an independent small-model build in the "fits on a floppy disk"
range β not a fine-tune of a bigger model.
## Architecture
LLaMA-style decoder, built from scratch:
| Field | Value |
|---|---|
| Parameters | **4,912,992** |
| Layers | 6 |
| Hidden size | 224 |
| Attention | GQA β 7 query heads, 2 KV heads, head_dim 32 |
| FFN | SwiGLU, intermediate 576 |
| Norm | RMSNorm (eps 1e-6) |
| Positional | RoPE (theta 10000) |
| Embeddings | Tied (input = output head) |
| Vocab | 8192 (BPE, trained on the corpus) |
| Context | 512 |
| Dtype | float32 |
## Training
- **Data:** `/corpus_slice300m` β 300 MB of diverse real web text,
tokenized to ~76.06M tokens (BPE, 8192 vocab).
- **Steps:** 10,000 @ batch 32 Γ seq 512
- **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1
- **LR:** 3e-4, cosine decay with 10% warmup, floor 10%
- **Grad clip:** 1.0
- **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA
## Measured results (independently recomputed)
- **Val perplexity: 58.91** β computed on the 1.52M-token held-out tail
(last 2% of the corpus), token-level, by the author.
- **Unigram baseline: 1453.67** on the same held-out tail.
- The model beats the unigram floor by ~25Γ, i.e. it genuinely learned
context, not just token frequencies.
Note: the in-training `val_loss` (β0.008) is **not** a reliable number β the
training loop's validation slice leaked from the training stream. The 58.91
above is the honest held-out figure.
## What it is good at / not
At 5M parameters this model produces grammatical first sentences but is far
from fluent: greedy decoding is coherent for roughly the first sentence and
then collapses into repetition loops (e.g. "the world's largest city in the
world is the world's largest cityβ¦"), and sampled decoding (top-p 0.9,
temp 0.7) is more varied but still drifts into incoherence within a few
sentences. Factual recall is weak. It is a demonstration of from-scratch
small-model training, not a useful general assistant.
### Sample outputs (greedy, temp 0.0)
> **The capital of France is** β "The capital of France is a very important
> part of the world's economy."
> **In machine learning, a neural network** β "In machine learning, a neural
> network is a very important part of the development of the system."
> **To make a cup of tea, you need** β "To make a cup of tea, you need to be
> sure to use a cup of coffee."
### Sample outputs (temp 0.7, top-p 0.9)
> **Once upon a time, there was a** β "Once upon a time, there was a great
> chance to have a lot of life."
> **The sun rises in the** β "The sun rises in the sun. Collecting a land,
> which is known for its brightness in a space, has been caused by the dirt
> and swords."
## Files
## Version
This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.
- `model.safetensors` β weights (27 MB, F32). `head.weight` and `tok.weight`
are tied (identical values); both keys are present for loaders that
expect an untied head.
- `tokenizer.json` β BPE tokenizer (8192 vocab)
- `config.json` β architecture config
## Reproduction
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
script, tokenizer and training log are not bundled here; the architecture is
fully specified in `config.json` and above. |