Add model card: CompactLM-5M from-scratch 4.9M-param GQA LLaMA
Browse files
README.md
CHANGED
|
@@ -2,130 +2,103 @@
|
|
| 2 |
license: apache-2.0
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
language: en
|
| 5 |
-
datasets:
|
| 6 |
-
- HuggingFaceFW/fineweb-edu
|
| 7 |
tags:
|
| 8 |
-
- tiny
|
| 9 |
- tiny-lm
|
| 10 |
-
- tiny
|
| 11 |
- slm
|
| 12 |
- small-language-model
|
|
|
|
| 13 |
- from-scratch
|
| 14 |
-
-
|
| 15 |
metrics:
|
| 16 |
- perplexity
|
| 17 |
---
|
| 18 |
|
| 19 |
# CompactLM-5M
|
| 20 |
|
| 21 |
-
A ~
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
A small causal language model in the spirit of the original LLaMA, trained
|
| 27 |
-
from scratch on an educational text corpus. It is a research/teaching artifact
|
| 28 |
-
showing what a clean, minimal transformer can do at the ~6M scale.
|
| 29 |
|
| 30 |
## Architecture
|
| 31 |
|
| 32 |
-
|
|
|
|
|
|
|
| 33 |
|---|---|
|
| 34 |
-
| Parameters | **
|
| 35 |
-
| Layers |
|
| 36 |
-
|
|
| 37 |
-
|
|
| 38 |
-
| FFN
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
|
| 40 |
| Context | 512 |
|
| 41 |
-
| Norm | RMSNorm, pre-norm |
|
| 42 |
-
| Attention | causal, RoPE (base 10000) |
|
| 43 |
-
| Embeddings | tied (`tok.weight` == `head.weight`) |
|
| 44 |
| Dtype | float32 |
|
| 45 |
|
| 46 |
-
Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
|
| 47 |
-
`RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.
|
| 48 |
-
|
| 49 |
## Training
|
| 50 |
|
| 51 |
-
- **Data:**
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
- **
|
| 55 |
-
- **
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
- **
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
>
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
|
| 99 |
## Files
|
| 100 |
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
`CompactLM` class from `train_compactlm5m.py`:
|
| 113 |
-
|
| 114 |
-
```python
|
| 115 |
-
import sys, torch
|
| 116 |
-
sys.path.insert(0, "<path-to-this-repo>")
|
| 117 |
-
from train_compactlm5m import CompactLM
|
| 118 |
-
from tokenizers import Tokenizer
|
| 119 |
-
|
| 120 |
-
tok = Tokenizer.from_file("tokenizer.json")
|
| 121 |
-
m = CompactLM(12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512).eval()
|
| 122 |
-
|
| 123 |
-
from safetensors.torch import load_file
|
| 124 |
-
sd = {k: v for k, v in load_file("model.safetensors").items()
|
| 125 |
-
if not k.startswith("head.weight")} # head.weight is tied to tok.weight
|
| 126 |
-
m.load_state_dict(sd, strict=False)
|
| 127 |
-
m.head.weight = m.tok.weight
|
| 128 |
-
|
| 129 |
-
ids = tok.encode("The cat sat on the").ids
|
| 130 |
-
# ... run m.forward on ids, sample, decode
|
| 131 |
-
```
|
|
|
|
| 2 |
license: apache-2.0
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
language: en
|
|
|
|
|
|
|
| 5 |
tags:
|
|
|
|
| 6 |
- tiny-lm
|
| 7 |
+
- tiny
|
| 8 |
- slm
|
| 9 |
- small-language-model
|
| 10 |
+
- sub-1m
|
| 11 |
- from-scratch
|
| 12 |
+
- text-generation
|
| 13 |
metrics:
|
| 14 |
- perplexity
|
| 15 |
---
|
| 16 |
|
| 17 |
# CompactLM-5M
|
| 18 |
|
| 19 |
+
A **from-scratch** ~5M-parameter language model, trained from zero on a
|
| 20 |
+
300 MB slice of diverse real web text (chemistry, code, literature, general
|
| 21 |
+
web). This is an independent small-model build in the "fits on a floppy disk"
|
| 22 |
+
range β not a fine-tune of a bigger model.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
|
| 24 |
## Architecture
|
| 25 |
|
| 26 |
+
LLaMA-style decoder, built from scratch:
|
| 27 |
+
|
| 28 |
+
| Field | Value |
|
| 29 |
|---|---|
|
| 30 |
+
| Parameters | **4,912,992** |
|
| 31 |
+
| Layers | 6 |
|
| 32 |
+
| Hidden size | 224 |
|
| 33 |
+
| Attention | GQA β 7 query heads, 2 KV heads, head_dim 32 |
|
| 34 |
+
| FFN | SwiGLU, intermediate 576 |
|
| 35 |
+
| Norm | RMSNorm (eps 1e-6) |
|
| 36 |
+
| Positional | RoPE (theta 10000) |
|
| 37 |
+
| Embeddings | Tied (input = output head) |
|
| 38 |
+
| Vocab | 8192 (BPE, trained on the corpus) |
|
| 39 |
| Context | 512 |
|
|
|
|
|
|
|
|
|
|
| 40 |
| Dtype | float32 |
|
| 41 |
|
|
|
|
|
|
|
|
|
|
| 42 |
## Training
|
| 43 |
|
| 44 |
+
- **Data:** `/corpus_slice300m` β 300 MB of diverse real web text,
|
| 45 |
+
tokenized to ~76.06M tokens (BPE, 8192 vocab).
|
| 46 |
+
- **Steps:** 10,000 @ batch 32 Γ seq 512
|
| 47 |
+
- **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1
|
| 48 |
+
- **LR:** 3e-4, cosine decay with 10% warmup, floor 10%
|
| 49 |
+
- **Grad clip:** 1.0
|
| 50 |
+
- **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA
|
| 51 |
+
|
| 52 |
+
## Measured results (independently recomputed)
|
| 53 |
+
|
| 54 |
+
- **Val perplexity: 58.91** β computed on the 1.52M-token held-out tail
|
| 55 |
+
(last 2% of the corpus), token-level, by the author.
|
| 56 |
+
- **Unigram baseline: 1453.67** on the same held-out tail.
|
| 57 |
+
- The model beats the unigram floor by ~25Γ, i.e. it genuinely learned
|
| 58 |
+
context, not just token frequencies.
|
| 59 |
+
|
| 60 |
+
Note: the in-training `val_loss` (β0.008) is **not** a reliable number β the
|
| 61 |
+
training loop's validation slice leaked from the training stream. The 58.91
|
| 62 |
+
above is the honest held-out figure.
|
| 63 |
+
|
| 64 |
+
## What it is good at / not
|
| 65 |
+
|
| 66 |
+
At 5M parameters this model produces fluent, grammatical, on-topic English,
|
| 67 |
+
but it is a small model: factual recall is weak and greedy decoding drifts
|
| 68 |
+
into repetition. Sampled decoding (top-p 0.9, temp 0.7) is noticeably more
|
| 69 |
+
diverse. It is a demonstration of from-scratch small-model training, not a
|
| 70 |
+
useful general assistant.
|
| 71 |
+
|
| 72 |
+
### Sample outputs (greedy, temp 0.0)
|
| 73 |
+
|
| 74 |
+
> **The capital of France is** β "The capital of France is a very important
|
| 75 |
+
> part of the world's economy."
|
| 76 |
+
|
| 77 |
+
> **In machine learning, a neural network** β "In machine learning, a neural
|
| 78 |
+
> network is a very important part of the development of the system."
|
| 79 |
+
|
| 80 |
+
> **To make a cup of tea, you need** β "To make a cup of tea, you need to be
|
| 81 |
+
> sure to use a cup of coffee."
|
| 82 |
+
|
| 83 |
+
### Sample outputs (temp 0.7, top-p 0.9)
|
| 84 |
+
|
| 85 |
+
> **Once upon a time, there was a** β "Once upon a time, there was a great
|
| 86 |
+
> chance to have a lot of life."
|
| 87 |
+
|
| 88 |
+
> **The sun rises in the** β "The sun rises in the sun. Collecting a land,
|
| 89 |
+
> which is known for its brightness in a space, has been caused by the dirt
|
| 90 |
+
> and swords."
|
| 91 |
|
| 92 |
## Files
|
| 93 |
|
| 94 |
+
- `model.safetensors` β weights (27 MB, F32). `head.weight` and `tok.weight`
|
| 95 |
+
are tied (identical values); both keys are present for loaders that
|
| 96 |
+
expect an untied head.
|
| 97 |
+
- `tokenizer.json` β BPE tokenizer (8192 vocab)
|
| 98 |
+
- `config.json` β architecture config
|
| 99 |
+
|
| 100 |
+
## Reproduction
|
| 101 |
+
|
| 102 |
+
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
|
| 103 |
+
script, tokenizer and training log are not bundled here; the architecture is
|
| 104 |
+
fully specified in `config.json` and above.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|