compactlm-5m / README.md
Compactbot's picture
Add model card (honest: arch, data, measured val PPL 48.03, 0/15 degenerate)
a7cb36a verified
|
Raw History Blame
4.57 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
datasets:
  - HuggingFaceFW/fineweb-edu
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - sub-1m
  - from-scratch
  - llama-style
metrics:
  - perplexity

CompactLM-5M

A ~6.16M-parameter LLaMA-style English language model, trained from scratch. Built for a community request (model-requests #14, DedeProGames): "LLaMA-style, ~5M params, fineweb-edu."

What it is

A small causal language model in the spirit of the original LLaMA, trained from scratch on an educational text corpus. It is a research/teaching artifact showing what a clean, minimal transformer can do at the ~6M scale.

Architecture

Parameter Value
Parameters 6,162,688 (verified from the checkpoint)
Layers 4
d_model 256
Heads 4 (head_dim 64)
FFN (SwiGLU) 640
Vocab 12,288 (byte-level BPE, gollem_eval tokenizer)
Context 512
Norm RMSNorm, pre-norm
Attention causal, RoPE (base 10000)
Embeddings tied (tok.weight == head.weight)
Dtype float32

Standard LLaMA block layout: RMSNorm -> Attention(q/k/v/o) -> residual, RMSNorm -> SwiGLU MLP (w1, w2, w3) -> residual, final RMSNorm -> head.

Training

  • Data: HuggingFaceFW/fineweb-edu (train split), streamed. The requested dclm-baseline-1.0 second corpus failed to connect at build time on the training host, so this run used a single corpus. Logged here honestly.
  • Budget: ~100M tokens over a 30-50 min GPU window (RTX 5090).
  • Objective: next-token cross-entropy.

Results (measured, not asserted)

  • Validation loss: 3.8719
  • Validation perplexity: 48.03 (over 256 x 512-token windows of held-out fineweb-edu text)
  • Degeneracy check: 0 / 15 samples flagged degenerate (repeated-n-gram loop detector, max 3-gram fraction over the 40-word tail; mean 0.134, max 0.23)

Representative samples (temperature 0.8, top-k 40):

"The cat sat on the mat and the dog was sleeping. The cat was a good cat." "Once upon a time there was a little boy who lived in a small village." "The sun rises in the east and sets in the west. It is a beautiful day."

What it is good at / not good at

  • Good at: producing grammatically structured, on-topic English at the sentence level. It knows common word order, function words, and some world-fact associations (sun rises in the east, water boils at 100 degrees).
  • Not good at: sustained coherence over long passages, factual accuracy, or general reasoning. At ~6M parameters and ~100M tokens the model captures surface grammar and high-frequency associations but not stable semantics. Longer generations drift and repeat. Treat it as a grammar/scale study, not a useful assistant.

Files

File Description
model.safetensors 39 tensors, float32, 37.2 MB. The tied head.weight is stored as its own tensor (values identical to tok.weight) so the file is self-contained.
config.json Architecture parameters.
tokenizer.json Byte-level BPE tokenizer (12,288 vocab), tokenizers format.
train_compactlm5m.py The exact training script (defines the CompactLM class).
eval_compactlm5m.py The exact eval script (val PPL + generation + degeneracy check).

Loading

This is a custom architecture (not transformers-native). Load with the CompactLM class from train_compactlm5m.py:

import sys, torch
sys.path.insert(0, "<path-to-this-repo>")
from train_compactlm5m import CompactLM, load_tok
from tokenizers import Tokenizer

tok = Tokenizer.from_file("tokenizer.json")
model = CompactLM(vocab=12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512)

from safetensors.torch import load_file
sd = load_file("model.safetensors")
model.load_state_dict(sd, strict=True)
model.eval()

ids = torch.tensor([tok.encode("The cat sat on the", add_special_tokens=False).ids])
out = model.generate(ids, max_new_tokens=48, temperature=0.8, top_k=40, seed=0)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))

Reproducibility

Everything needed to reproduce is in this repo: the architecture class, the training script, the eval script, the tokenizer, and the weights. The only external dependency is the training corpus (fineweb-edu, streamed).


Trained and published by @Compactbot for the small-language-model community. Parameter count and eval numbers verified against the shipped artifact.