| --- |
| license: apache-2.0 |
| pipeline_tag: text-generation |
| language: en |
| tags: |
| - tiny-lm |
| - tiny |
| - slm |
| - small-language-model |
| - sub-1m |
| - from-scratch |
| - text-generation |
| metrics: |
| - perplexity |
| --- |
| |
| # CompactLM-5M |
|
|
| A **from-scratch** ~5M-parameter language model, trained from zero on a |
| 300 MB slice of diverse real web text (chemistry, code, literature, general |
| web). This is an independent small-model build in the "fits on a floppy disk" |
| range β not a fine-tune of a bigger model. |
|
|
| ## Architecture |
|
|
| LLaMA-style decoder, built from scratch: |
|
|
| | Field | Value | |
| |---|---| |
| | Parameters | **4,912,992** | |
| | Layers | 6 | |
| | Hidden size | 224 | |
| | Attention | GQA β 7 query heads, 2 KV heads, head_dim 32 | |
| | FFN | SwiGLU, intermediate 576 | |
| | Norm | RMSNorm (eps 1e-6) | |
| | Positional | RoPE (theta 10000) | |
| | Embeddings | Tied (input = output head) | |
| | Vocab | 8192 (BPE, trained on the corpus) | |
| | Context | 512 | |
| | Dtype | float32 | |
| |
| ## Training |
| |
| - **Data:** `/corpus_slice300m` β 300 MB of diverse real web text, |
| tokenized to ~76.06M tokens (BPE, 8192 vocab). |
| - **Steps:** 10,000 @ batch 32 Γ seq 512 |
| - **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1 |
| - **LR:** 3e-4, cosine decay with 10% warmup, floor 10% |
| - **Grad clip:** 1.0 |
| - **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA |
|
|
| ## Measured results (independently recomputed) |
|
|
| - **Val perplexity: 58.91** β computed on the 1.52M-token held-out tail |
| (last 2% of the corpus), token-level, by the author. |
| - **Unigram baseline: 1453.67** on the same held-out tail. |
| - The model beats the unigram floor by ~25Γ, i.e. it genuinely learned |
| context, not just token frequencies. |
|
|
| Note: the in-training `val_loss` (β0.008) is **not** a reliable number β the |
| training loop's validation slice leaked from the training stream. The 58.91 |
| above is the honest held-out figure. |
|
|
| ## What it is good at / not |
|
|
| At 5M parameters this model produces grammatical first sentences but is far |
| from fluent: greedy decoding is coherent for roughly the first sentence and |
| then collapses into repetition loops (e.g. "the world's largest city in the |
| world is the world's largest cityβ¦"), and sampled decoding (top-p 0.9, |
| temp 0.7) is more varied but still drifts into incoherence within a few |
| sentences. Factual recall is weak. It is a demonstration of from-scratch |
| small-model training, not a useful general assistant. |
|
|
| ### Sample outputs (greedy, temp 0.0) |
|
|
| > **The capital of France is** β "The capital of France is a very important |
| > part of the world's economy." |
|
|
| > **In machine learning, a neural network** β "In machine learning, a neural |
| > network is a very important part of the development of the system." |
|
|
| > **To make a cup of tea, you need** β "To make a cup of tea, you need to be |
| > sure to use a cup of coffee." |
|
|
| ### Sample outputs (temp 0.7, top-p 0.9) |
|
|
| > **Once upon a time, there was a** β "Once upon a time, there was a great |
| > chance to have a lot of life." |
|
|
| > **The sun rises in the** β "The sun rises in the sun. Collecting a land, |
| > which is known for its brightness in a space, has been caused by the dirt |
| > and swords." |
|
|
| ## Files |
|
|
| ## Version |
|
|
| This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator. |
|
|
|
|
| - `model.safetensors` β weights (27 MB, F32). `head.weight` and `tok.weight` |
| are tied (identical values); both keys are present for loaders that |
| expect an untied head. |
| - `tokenizer.json` β BPE tokenizer (8192 vocab) |
| - `config.json` β architecture config |
|
|
| ## Reproduction |
|
|
| Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact |
| script, tokenizer and training log are not bundled here; the architecture is |
| fully specified in `config.json` and above. |