File size: 4,036 Bytes
da4f145
 
 
cafd02f
da4f145
 
a6ec395
da4f145
 
a6ec395
da4f145
a6ec395
da4f145
 
 
 
 
 
a6ec395
 
 
 
da4f145
 
 
a6ec395
 
 
da4f145
a6ec395
 
 
 
 
 
 
 
 
da4f145
cafd02f
da4f145
 
 
a6ec395
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9323e39
 
 
 
 
 
 
a6ec395
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
da4f145
 
 
5dbd13e
 
9323e39
5dbd13e
 
a6ec395
 
 
 
 
 
 
 
 
 
9323e39
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny-lm
  - tiny
  - slm
  - small-language-model
  - sub-1m
  - from-scratch
  - text-generation
metrics:
  - perplexity
---

# CompactLM-5M

A **from-scratch** ~5M-parameter language model, trained from zero on a
300 MB slice of diverse real web text (chemistry, code, literature, general
web). This is an independent small-model build in the "fits on a floppy disk"
range β€” not a fine-tune of a bigger model.

## Architecture

LLaMA-style decoder, built from scratch:

| Field | Value |
|---|---|
| Parameters | **4,912,992** |
| Layers | 6 |
| Hidden size | 224 |
| Attention | GQA β€” 7 query heads, 2 KV heads, head_dim 32 |
| FFN | SwiGLU, intermediate 576 |
| Norm | RMSNorm (eps 1e-6) |
| Positional | RoPE (theta 10000) |
| Embeddings | Tied (input = output head) |
| Vocab | 8192 (BPE, trained on the corpus) |
| Context | 512 |
| Dtype | float32 |

## Training

- **Data:** `/corpus_slice300m` β€” 300 MB of diverse real web text,
  tokenized to ~76.06M tokens (BPE, 8192 vocab).
- **Steps:** 10,000 @ batch 32 Γ— seq 512
- **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1
- **LR:** 3e-4, cosine decay with 10% warmup, floor 10%
- **Grad clip:** 1.0
- **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA

## Measured results (independently recomputed)

- **Val perplexity: 58.91** β€” computed on the 1.52M-token held-out tail
  (last 2% of the corpus), token-level, by the author.
- **Unigram baseline: 1453.67** on the same held-out tail.
- The model beats the unigram floor by ~25Γ—, i.e. it genuinely learned
  context, not just token frequencies.

Note: the in-training `val_loss` (β‰ˆ0.008) is **not** a reliable number β€” the
training loop's validation slice leaked from the training stream. The 58.91
above is the honest held-out figure.

## What it is good at / not

At 5M parameters this model produces grammatical first sentences but is far
from fluent: greedy decoding is coherent for roughly the first sentence and
then collapses into repetition loops (e.g. "the world's largest city in the
world is the world's largest city…"), and sampled decoding (top-p 0.9,
temp 0.7) is more varied but still drifts into incoherence within a few
sentences. Factual recall is weak. It is a demonstration of from-scratch
small-model training, not a useful general assistant.

### Sample outputs (greedy, temp 0.0)

> **The capital of France is** β†’ "The capital of France is a very important
> part of the world's economy."

> **In machine learning, a neural network** β†’ "In machine learning, a neural
> network is a very important part of the development of the system."

> **To make a cup of tea, you need** β†’ "To make a cup of tea, you need to be
> sure to use a cup of coffee."

### Sample outputs (temp 0.7, top-p 0.9)

> **Once upon a time, there was a** β†’ "Once upon a time, there was a great
> chance to have a lot of life."

> **The sun rises in the** β†’ "The sun rises in the sun. Collecting a land,
> which is known for its brightness in a space, has been caused by the dirt
> and swords."

## Files

## Version

This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.


- `model.safetensors` β€” weights (27 MB, F32). `head.weight` and `tok.weight`
  are tied (identical values); both keys are present for loaders that
  expect an untied head.
- `tokenizer.json` β€” BPE tokenizer (8192 vocab)
- `config.json` β€” architecture config

## Reproduction

Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
script, tokenizer and training log are not bundled here; the architecture is
fully specified in `config.json` and above.