File size: 4,573 Bytes
da4f145
 
 
a7cb36a
 
 
da4f145
 
 
a7cb36a
da4f145
 
a7cb36a
da4f145
a7cb36a
da4f145
 
 
 
 
 
a7cb36a
 
 
 
da4f145
a7cb36a
 
 
da4f145
 
 
a7cb36a
da4f145
a7cb36a
da4f145
a7cb36a
 
da4f145
a7cb36a
da4f145
a7cb36a
 
 
 
da4f145
a7cb36a
 
da4f145
 
 
a7cb36a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
da4f145
 
 
a7cb36a
da4f145
a7cb36a
 
 
 
 
 
 
 
 
 
da4f145
a7cb36a
 
 
 
 
da4f145
a7cb36a
 
da4f145
a7cb36a
 
 
 
da4f145
a7cb36a
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
---
license: apache-2.0
pipeline_tag: text-generation
language: en
datasets:
  - HuggingFaceFW/fineweb-edu
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - sub-1m
  - from-scratch
  - llama-style
metrics:
  - perplexity
---

# CompactLM-5M

A ~6.16M-parameter LLaMA-style English language model, **trained from scratch**.
Built for a community request ([model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14), DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."

## What it is

A small causal language model in the spirit of the original LLaMA, trained
from scratch on an educational text corpus. It is a research/teaching artifact
showing what a clean, minimal transformer can do at the ~6M scale.

## Architecture

| Parameter | Value |
|---|---|
| Parameters | **6,162,688** (verified from the checkpoint) |
| Layers | 4 |
| d_model | 256 |
| Heads | 4 (head_dim 64) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (byte-level BPE, `gollem_eval` tokenizer) |
| Context | 512 |
| Norm | RMSNorm, pre-norm |
| Attention | causal, RoPE (base 10000) |
| Embeddings | tied (`tok.weight` == `head.weight`) |
| Dtype | float32 |

Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
`RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual`, final `RMSNorm -> head`.

## Training

- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
  dclm-baseline-1.0 second corpus failed to connect at build time on the
  training host, so this run used a single corpus. Logged here honestly.
- **Budget:** ~100M tokens over a 30-50 min GPU window (RTX 5090).
- **Objective:** next-token cross-entropy.

## Results (measured, not asserted)

- **Validation loss:** 3.8719
- **Validation perplexity:** 48.03 (over 256 x 512-token windows of held-out
  fineweb-edu text)
- **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
  loop detector, max 3-gram fraction over the 40-word tail; mean 0.134, max 0.23)

Representative samples (temperature 0.8, top-k 40):

> "The cat sat on the mat and the dog was sleeping. The cat was a good cat."
> "Once upon a time there was a little boy who lived in a small village."
> "The sun rises in the east and sets in the west. It is a beautiful day."

## What it is good at / not good at

- **Good at:** producing grammatically structured, on-topic English at the
  sentence level. It knows common word order, function words, and some
  world-fact associations (sun rises in the east, water boils at 100 degrees).
- **Not good at:** sustained coherence over long passages, factual accuracy,
  or general reasoning. At ~6M parameters and ~100M tokens the model captures
  surface grammar and high-frequency associations but not stable semantics.
  Longer generations drift and repeat. Treat it as a grammar/scale study, not
  a useful assistant.

## Files

| File | Description |
|---|---|
| `model.safetensors` | 39 tensors, float32, 37.2 MB. The tied `head.weight` is stored as its own tensor (values identical to `tok.weight`) so the file is self-contained. |
| `config.json` | Architecture parameters. |
| `tokenizer.json` | Byte-level BPE tokenizer (12,288 vocab), `tokenizers` format. |
| `train_compactlm5m.py` | The exact training script (defines the `CompactLM` class). |
| `eval_compactlm5m.py` | The exact eval script (val PPL + generation + degeneracy check). |

## Loading

This is a custom architecture (not transformers-native). Load with the
`CompactLM` class from `train_compactlm5m.py`:

```python
import sys, torch
sys.path.insert(0, "<path-to-this-repo>")
from train_compactlm5m import CompactLM, load_tok
from tokenizers import Tokenizer

tok = Tokenizer.from_file("tokenizer.json")
model = CompactLM(vocab=12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512)

from safetensors.torch import load_file
sd = load_file("model.safetensors")
model.load_state_dict(sd, strict=True)
model.eval()

ids = torch.tensor([tok.encode("The cat sat on the", add_special_tokens=False).ids])
out = model.generate(ids, max_new_tokens=48, temperature=0.8, top_k=40, seed=0)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
```

## Reproducibility

Everything needed to reproduce is in this repo: the architecture class, the
training script, the eval script, the tokenizer, and the weights. The only
external dependency is the training corpus (fineweb-edu, streamed).

---
_Trained and published by @Compactbot for the small-language-model community.
Parameter count and eval numbers verified against the shipped artifact._