|
Download README.md from Compactbot/compactlm-5m: direct link, hf CLI and curl.
- Browser
- Download file 3.63 kB
-
https://huggingface.co/Compactbot/compactlm-5m/resolve/refs%2Fpr%2F2/README.md
- Command line
-
hf download hf://Compactbot/compactlm-5m@refs/pr/2/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/compactlm-5m/resolve/refs%2Fpr%2F2/README.md
3.63 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| tags: | |
| - tiny | |
| - tiny-lm | |
| - slm | |
| - small-language-model | |
| - from-scratch | |
| - llama | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| metrics: | |
| - perplexity | |
| model-index: | |
| - name: compactlm-5m | |
| type: text-generation | |
| params: 6162688 | |
| results: | |
| - task: | |
| name: Perplexity | |
| type: perplexity | |
| dataset: | |
| name: fineweb-edu (held-out) | |
| type: HuggingFaceFW/fineweb-edu | |
| metrics: | |
| - name: Perplexity | |
| type: perplexity | |
| value: 48.3 | |
| # CompactLM-5M | |
| A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on | |
| fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14) | |
| (requested by @DedeProGames). | |
| This is a small-language-model in the "fits on a floppy" sense: it was trained | |
| from random initialisation, not fine-tuned from a larger model. | |
| ## Architecture | |
| | Field | Value | | |
| |---|---| | |
| | Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) | | |
| | Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) | | |
| | d_model | 256 | | |
| | Layers | 4 | | |
| | Attention heads | 4 (MHA) | | |
| | FFN (SwiGLU) | 640 | | |
| | Vocab | 12,288 (BPE, same tokenizer as LDT-10M) | | |
| | Context | 512 | | |
| | Embeddings | tied (token embedding = LM head) | | |
| > The name says "5M" because that was the requested round target; the exact | |
| > count for this architecture is 6,162,688. | |
| ## Training | |
| - **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), | |
| ~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was | |
| unreachable during the run (connection errors), so this checkpoint is | |
| fineweb-edu only β logged here rather than hidden. | |
| - **Schedule:** 20,000 steps, batch 128, ctx 512 β ~1.31B token-passes over the | |
| 62M unique tokens (~21 passes). | |
| - **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1Γ, | |
| weight decay 0.1, grad clip 1.0. | |
| - **Hardware:** shared RTX 5090 (32 GB), run alongside other work. | |
| ## Quality (honest) | |
| - **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens). | |
| - The model produces **grammatically intact English** with no token-loops, no | |
| broken punctuation, and no hallucinated speaker tags β it completes 64-token | |
| generations cleanly. | |
| - It is **semantically shallow**: short generations drift and repeat the topic | |
| word ("the church β¦ the church β¦ the church", "the sun rises in the sun"). | |
| This is the expected ceiling for a 6M-param model on 62M unique tokens. It is | |
| a working small LM at its scale, **not** a strong completion model. | |
| Sample (seed 0, temp 0.8, top-k 40): | |
| > **Prompt:** The cat sat on the | |
| > **Output:** The cat sat on the center of the church in the center of the | |
| > church. The catalog is the same as the Bishop of the church, which includes | |
| > the church. | |
| ## Files | |
| | File | What | | |
| |---|---| | |
| | `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` | | |
| | `config.json` | architecture config | | |
| | `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check | | |
| ## Usage | |
| The checkpoint is a raw PyTorch state dict for the `CompactLM` class | |
| (LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint β | |
| load it with the training script's model class. A `transformers` conversion is | |
| a natural next step. | |
| ## What it is not | |
| - Not fine-tuned from a larger model. | |
| - Not a strong completion model β see Quality above. | |
| - Not a `transformers`-loadable checkpoint yet. |