compactlm-5m / README.md
Compactbot's picture
Add CompactLM-5M: from-scratch LLaMA-style 6.2M-param English LM on fineweb-edu
da4f145 verified
|
Raw History Blame
3.63 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- tiny
- tiny-lm
- slm
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
metrics:
- perplexity
model-index:
- name: compactlm-5m
type: text-generation
params: 6162688
results:
- task:
name: Perplexity
type: perplexity
dataset:
name: fineweb-edu (held-out)
type: HuggingFaceFW/fineweb-edu
metrics:
- name: Perplexity
type: perplexity
value: 48.3
---
# CompactLM-5M
A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on
fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14)
(requested by @DedeProGames).
This is a small-language-model in the "fits on a floppy" sense: it was trained
from random initialisation, not fine-tuned from a larger model.
## Architecture
| Field | Value |
|---|---|
| Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) |
| Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) |
| d_model | 256 |
| Layers | 4 |
| Attention heads | 4 (MHA) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (BPE, same tokenizer as LDT-10M) |
| Context | 512 |
| Embeddings | tied (token embedding = LM head) |
> The name says "5M" because that was the requested round target; the exact
> count for this architecture is 6,162,688.
## Training
- **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was
unreachable during the run (connection errors), so this checkpoint is
fineweb-edu only β€” logged here rather than hidden.
- **Schedule:** 20,000 steps, batch 128, ctx 512 β†’ ~1.31B token-passes over the
62M unique tokens (~21 passes).
- **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1Γ—,
weight decay 0.1, grad clip 1.0.
- **Hardware:** shared RTX 5090 (32 GB), run alongside other work.
## Quality (honest)
- **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens).
- The model produces **grammatically intact English** with no token-loops, no
broken punctuation, and no hallucinated speaker tags β€” it completes 64-token
generations cleanly.
- It is **semantically shallow**: short generations drift and repeat the topic
word ("the church … the church … the church", "the sun rises in the sun").
This is the expected ceiling for a 6M-param model on 62M unique tokens. It is
a working small LM at its scale, **not** a strong completion model.
Sample (seed 0, temp 0.8, top-k 40):
> **Prompt:** The cat sat on the
> **Output:** The cat sat on the center of the church in the center of the
> church. The catalog is the same as the Bishop of the church, which includes
> the church.
## Files
| File | What |
|---|---|
| `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` |
| `config.json` | architecture config |
| `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check |
## Usage
The checkpoint is a raw PyTorch state dict for the `CompactLM` class
(LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint β€”
load it with the training script's model class. A `transformers` conversion is
a natural next step.
## What it is not
- Not fine-tuned from a larger model.
- Not a strong completion model β€” see Quality above.
- Not a `transformers`-loadable checkpoint yet.