File size: 3,629 Bytes
da4f145
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
---
license: apache-2.0
language:
  - en
pipeline_tag: text-generation
library_name: transformers
tags:
  - tiny
  - tiny-lm
  - slm
  - small-language-model
  - from-scratch
  - llama
datasets:
  - HuggingFaceFW/fineweb-edu
metrics:
  - perplexity
model-index:
  - name: compactlm-5m
    type: text-generation
    params: 6162688
    results:
      - task:
          name: Perplexity
          type: perplexity
        dataset:
          name: fineweb-edu (held-out)
          type: HuggingFaceFW/fineweb-edu
        metrics:
          - name: Perplexity
            type: perplexity
            value: 48.3
---

# CompactLM-5M

A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on
fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14)
(requested by @DedeProGames).

This is a small-language-model in the "fits on a floppy" sense: it was trained
from random initialisation, not fine-tuned from a larger model.

## Architecture

| Field | Value |
|---|---|
| Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) |
| Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) |
| d_model | 256 |
| Layers | 4 |
| Attention heads | 4 (MHA) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (BPE, same tokenizer as LDT-10M) |
| Context | 512 |
| Embeddings | tied (token embedding = LM head) |

> The name says "5M" because that was the requested round target; the exact
> count for this architecture is 6,162,688.

## Training

- **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
  ~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was
  unreachable during the run (connection errors), so this checkpoint is
  fineweb-edu only — logged here rather than hidden.
- **Schedule:** 20,000 steps, batch 128, ctx 512 → ~1.31B token-passes over the
  62M unique tokens (~21 passes).
- **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×,
  weight decay 0.1, grad clip 1.0.
- **Hardware:** shared RTX 5090 (32 GB), run alongside other work.

## Quality (honest)

- **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens).
- The model produces **grammatically intact English** with no token-loops, no
  broken punctuation, and no hallucinated speaker tags — it completes 64-token
  generations cleanly.
- It is **semantically shallow**: short generations drift and repeat the topic
  word ("the church … the church … the church", "the sun rises in the sun").
  This is the expected ceiling for a 6M-param model on 62M unique tokens. It is
  a working small LM at its scale, **not** a strong completion model.

Sample (seed 0, temp 0.8, top-k 40):

> **Prompt:** The cat sat on the
> **Output:** The cat sat on the center of the church in the center of the
> church. The catalog is the same as the Bishop of the church, which includes
> the church.

## Files

| File | What |
|---|---|
| `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` |
| `config.json` | architecture config |
| `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check |

## Usage

The checkpoint is a raw PyTorch state dict for the `CompactLM` class
(LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint —
load it with the training script's model class. A `transformers` conversion is
a natural next step.

## What it is not

- Not fine-tuned from a larger model.
- Not a strong completion model — see Quality above.
- Not a `transformers`-loadable checkpoint yet.