File size: 3,629 Bytes
da4f145 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 | ---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- tiny
- tiny-lm
- slm
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
metrics:
- perplexity
model-index:
- name: compactlm-5m
type: text-generation
params: 6162688
results:
- task:
name: Perplexity
type: perplexity
dataset:
name: fineweb-edu (held-out)
type: HuggingFaceFW/fineweb-edu
metrics:
- name: Perplexity
type: perplexity
value: 48.3
---
# CompactLM-5M
A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on
fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14)
(requested by @DedeProGames).
This is a small-language-model in the "fits on a floppy" sense: it was trained
from random initialisation, not fine-tuned from a larger model.
## Architecture
| Field | Value |
|---|---|
| Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) |
| Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) |
| d_model | 256 |
| Layers | 4 |
| Attention heads | 4 (MHA) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (BPE, same tokenizer as LDT-10M) |
| Context | 512 |
| Embeddings | tied (token embedding = LM head) |
> The name says "5M" because that was the requested round target; the exact
> count for this architecture is 6,162,688.
## Training
- **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was
unreachable during the run (connection errors), so this checkpoint is
fineweb-edu only — logged here rather than hidden.
- **Schedule:** 20,000 steps, batch 128, ctx 512 → ~1.31B token-passes over the
62M unique tokens (~21 passes).
- **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×,
weight decay 0.1, grad clip 1.0.
- **Hardware:** shared RTX 5090 (32 GB), run alongside other work.
## Quality (honest)
- **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens).
- The model produces **grammatically intact English** with no token-loops, no
broken punctuation, and no hallucinated speaker tags — it completes 64-token
generations cleanly.
- It is **semantically shallow**: short generations drift and repeat the topic
word ("the church … the church … the church", "the sun rises in the sun").
This is the expected ceiling for a 6M-param model on 62M unique tokens. It is
a working small LM at its scale, **not** a strong completion model.
Sample (seed 0, temp 0.8, top-k 40):
> **Prompt:** The cat sat on the
> **Output:** The cat sat on the center of the church in the center of the
> church. The catalog is the same as the Bishop of the church, which includes
> the church.
## Files
| File | What |
|---|---|
| `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` |
| `config.json` | architecture config |
| `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check |
## Usage
The checkpoint is a raw PyTorch state dict for the `CompactLM` class
(LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint —
load it with the training script's model class. A `transformers` conversion is
a natural next step.
## What it is not
- Not fine-tuned from a larger model.
- Not a strong completion model — see Quality above.
- Not a `transformers`-loadable checkpoint yet. |