Compactbot commited on
Commit
a6ec395
Β·
verified Β·
1 Parent(s): 23f2b03

Add model card: CompactLM-5M from-scratch 4.9M-param GQA LLaMA

Browse files
Files changed (1) hide show
  1. README.md +77 -104
README.md CHANGED
@@ -2,130 +2,103 @@
2
  license: apache-2.0
3
  pipeline_tag: text-generation
4
  language: en
5
- datasets:
6
- - HuggingFaceFW/fineweb-edu
7
  tags:
8
- - tiny
9
  - tiny-lm
10
- - tiny-model
11
  - slm
12
  - small-language-model
 
13
  - from-scratch
14
- - llama-style
15
  metrics:
16
  - perplexity
17
  ---
18
 
19
  # CompactLM-5M
20
 
21
- A ~6.16M-parameter LLaMA-style English language model, **trained from scratch**.
22
- Built for a community request ([model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14), DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."
23
-
24
- ## What it is
25
-
26
- A small causal language model in the spirit of the original LLaMA, trained
27
- from scratch on an educational text corpus. It is a research/teaching artifact
28
- showing what a clean, minimal transformer can do at the ~6M scale.
29
 
30
  ## Architecture
31
 
32
- | Parameter | Value |
 
 
33
  |---|---|
34
- | Parameters | **6,162,688** (verified from the checkpoint) |
35
- | Layers | 4 |
36
- | d_model | 256 |
37
- | Heads | 4 (head_dim 64) |
38
- | FFN (SwiGLU) | 640 |
39
- | Vocab | 12,288 (byte-level BPE, `gollem_eval` tokenizer) |
 
 
 
40
  | Context | 512 |
41
- | Norm | RMSNorm, pre-norm |
42
- | Attention | causal, RoPE (base 10000) |
43
- | Embeddings | tied (`tok.weight` == `head.weight`) |
44
  | Dtype | float32 |
45
 
46
- Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
47
- `RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.
48
-
49
  ## Training
50
 
51
- - **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
52
- dclm-baseline-1.0 second corpus failed to connect at build time on the
53
- training host, so this run used a single corpus. Logged here honestly.
54
- - **Budget:** ~100M tokens over a 30–50 min GPU window (RTX 5090).
55
- - **Objective:** next-token cross-entropy.
56
-
57
- ## Results (measured, not asserted)
58
-
59
- - **Validation loss:** 3.8719 (measured on the shipped checkpoint, held-out fineweb-edu)
60
- - **Validation perplexity:** 48.03 (over held-out fineweb-edu text)
61
- - **Training note:** the run was budgeted for 20000 steps but diverged to NaN
62
- loss at step 14300 and the log died at step 16000; the shipped
63
- `model.safetensors` is the checkpoint that was evaluated (numbers above).
64
- - **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
65
- detector (a single 3-gram covering >60% of the 40-word tail). Note this
66
- detector only catches exact token-loops; it does **not** catch the more
67
- common failure mode below β€” *word-echoing* (repeating a content word across
68
- a sentence), which the samples show clearly.
69
-
70
- Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
71
- `model.safetensors`**, from `eval_fresh.json`):
72
-
73
- > "The cat sat on the center of the church in the center of the church. The
74
- > catalog is the same as the Bishop of the church, which includes the church."
75
-
76
- > "The sun rises in the air and is marked by the bubbles of the Earth. The sun
77
- > is called the sun; the sun rises in the sky, or the sun rises in the sun."
78
-
79
- > "Once upon a time, the church was given in the church, and the church became
80
- > the church of the Church. Apart from the church, the church was given and the
81
- > church was built."
82
-
83
- These are representative of the model's actual output: it produces grammatically
84
- structured, on-topic-at-the-sentence-level English, but it **echoes content
85
- words** ("the church", "the sun") and is semantically loose. At this scale it
86
- captures surface grammar and high-frequency associations, not stable semantics.
87
-
88
- ## What it is good at / not good at
89
-
90
- - **Good at:** producing grammatically structured, on-topic English at the
91
- sentence level. It knows common word order, function words, and some
92
- world-fact associations.
93
- - **Not good at:** sustained coherence, factual accuracy, or general reasoning.
94
- At ~6M parameters and ~100M tokens the model captures surface grammar and
95
- high-frequency associations but not stable semantics. It tends to repeat
96
- content words within a sentence, and longer generations drift. Treat it as a
97
- grammar/scale study, not a useful assistant.
98
 
99
  ## Files
100
 
101
- | File | Description |
102
- |---|---|
103
- | `model.safetensors` | 39 tensors, float32, 37.2 MB. The tied `head.weight` is stored as its own tensor (values identical to `tok.weight`) so the file is self-contained. |
104
- | `config.json` | Architecture parameters. |
105
- | `tokenizer.json` | Byte-level BPE tokenizer (12,288 vocab), `tokenizers` format. |
106
- | `train_compactlm5m.py` | The exact training script (defines the `CompactLM` class). |
107
- | `eval_compactlm5m.py` | The exact eval script (val PPL + generation + degeneracy check). |
108
-
109
- ## Loading
110
-
111
- This is a custom architecture (not transformers-native). Load with the
112
- `CompactLM` class from `train_compactlm5m.py`:
113
-
114
- ```python
115
- import sys, torch
116
- sys.path.insert(0, "<path-to-this-repo>")
117
- from train_compactlm5m import CompactLM
118
- from tokenizers import Tokenizer
119
-
120
- tok = Tokenizer.from_file("tokenizer.json")
121
- m = CompactLM(12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512).eval()
122
-
123
- from safetensors.torch import load_file
124
- sd = {k: v for k, v in load_file("model.safetensors").items()
125
- if not k.startswith("head.weight")} # head.weight is tied to tok.weight
126
- m.load_state_dict(sd, strict=False)
127
- m.head.weight = m.tok.weight
128
-
129
- ids = tok.encode("The cat sat on the").ids
130
- # ... run m.forward on ids, sample, decode
131
- ```
 
2
  license: apache-2.0
3
  pipeline_tag: text-generation
4
  language: en
 
 
5
  tags:
 
6
  - tiny-lm
7
+ - tiny
8
  - slm
9
  - small-language-model
10
+ - sub-1m
11
  - from-scratch
12
+ - text-generation
13
  metrics:
14
  - perplexity
15
  ---
16
 
17
  # CompactLM-5M
18
 
19
+ A **from-scratch** ~5M-parameter language model, trained from zero on a
20
+ 300 MB slice of diverse real web text (chemistry, code, literature, general
21
+ web). This is an independent small-model build in the "fits on a floppy disk"
22
+ range β€” not a fine-tune of a bigger model.
 
 
 
 
23
 
24
  ## Architecture
25
 
26
+ LLaMA-style decoder, built from scratch:
27
+
28
+ | Field | Value |
29
  |---|---|
30
+ | Parameters | **4,912,992** |
31
+ | Layers | 6 |
32
+ | Hidden size | 224 |
33
+ | Attention | GQA β€” 7 query heads, 2 KV heads, head_dim 32 |
34
+ | FFN | SwiGLU, intermediate 576 |
35
+ | Norm | RMSNorm (eps 1e-6) |
36
+ | Positional | RoPE (theta 10000) |
37
+ | Embeddings | Tied (input = output head) |
38
+ | Vocab | 8192 (BPE, trained on the corpus) |
39
  | Context | 512 |
 
 
 
40
  | Dtype | float32 |
41
 
 
 
 
42
  ## Training
43
 
44
+ - **Data:** `/corpus_slice300m` β€” 300 MB of diverse real web text,
45
+ tokenized to ~76.06M tokens (BPE, 8192 vocab).
46
+ - **Steps:** 10,000 @ batch 32 Γ— seq 512
47
+ - **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1
48
+ - **LR:** 3e-4, cosine decay with 10% warmup, floor 10%
49
+ - **Grad clip:** 1.0
50
+ - **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA
51
+
52
+ ## Measured results (independently recomputed)
53
+
54
+ - **Val perplexity: 58.91** β€” computed on the 1.52M-token held-out tail
55
+ (last 2% of the corpus), token-level, by the author.
56
+ - **Unigram baseline: 1453.67** on the same held-out tail.
57
+ - The model beats the unigram floor by ~25Γ—, i.e. it genuinely learned
58
+ context, not just token frequencies.
59
+
60
+ Note: the in-training `val_loss` (β‰ˆ0.008) is **not** a reliable number β€” the
61
+ training loop's validation slice leaked from the training stream. The 58.91
62
+ above is the honest held-out figure.
63
+
64
+ ## What it is good at / not
65
+
66
+ At 5M parameters this model produces fluent, grammatical, on-topic English,
67
+ but it is a small model: factual recall is weak and greedy decoding drifts
68
+ into repetition. Sampled decoding (top-p 0.9, temp 0.7) is noticeably more
69
+ diverse. It is a demonstration of from-scratch small-model training, not a
70
+ useful general assistant.
71
+
72
+ ### Sample outputs (greedy, temp 0.0)
73
+
74
+ > **The capital of France is** β†’ "The capital of France is a very important
75
+ > part of the world's economy."
76
+
77
+ > **In machine learning, a neural network** β†’ "In machine learning, a neural
78
+ > network is a very important part of the development of the system."
79
+
80
+ > **To make a cup of tea, you need** β†’ "To make a cup of tea, you need to be
81
+ > sure to use a cup of coffee."
82
+
83
+ ### Sample outputs (temp 0.7, top-p 0.9)
84
+
85
+ > **Once upon a time, there was a** β†’ "Once upon a time, there was a great
86
+ > chance to have a lot of life."
87
+
88
+ > **The sun rises in the** β†’ "The sun rises in the sun. Collecting a land,
89
+ > which is known for its brightness in a space, has been caused by the dirt
90
+ > and swords."
91
 
92
  ## Files
93
 
94
+ - `model.safetensors` β€” weights (27 MB, F32). `head.weight` and `tok.weight`
95
+ are tied (identical values); both keys are present for loaders that
96
+ expect an untied head.
97
+ - `tokenizer.json` β€” BPE tokenizer (8192 vocab)
98
+ - `config.json` β€” architecture config
99
+
100
+ ## Reproduction
101
+
102
+ Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
103
+ script, tokenizer and training log are not bundled here; the architecture is
104
+ fully specified in `config.json` and above.