Add model card (honest: real samples, final val 3.8775 / ppl 48.30).
Browse files
README.md
CHANGED
|
@@ -10,7 +10,6 @@ tags:
|
|
| 10 |
- tiny-model
|
| 11 |
- slm
|
| 12 |
- small-language-model
|
| 13 |
-
- sub-1m
|
| 14 |
- from-scratch
|
| 15 |
- llama-style
|
| 16 |
metrics:
|
|
@@ -52,42 +51,41 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
|
|
| 52 |
- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
|
| 53 |
dclm-baseline-1.0 second corpus failed to connect at build time on the
|
| 54 |
training host, so this run used a single corpus. Logged here honestly.
|
| 55 |
-
- **Budget:** ~100M tokens over a 30
|
| 56 |
- **Objective:** next-token cross-entropy.
|
| 57 |
|
| 58 |
## Results (measured, not asserted)
|
| 59 |
|
| 60 |
-
- **Validation loss:** 3.
|
| 61 |
-
- **Validation perplexity:** 48.
|
| 62 |
-
fineweb-edu text)
|
| 63 |
- **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
|
| 64 |
-
loop detector, max 3-gram fraction over the 40-word tail
|
| 65 |
|
| 66 |
-
Representative samples (temperature 0.8, top-k 40,
|
| 67 |
-
|
| 68 |
|
| 69 |
-
> "The cat sat on the
|
| 70 |
-
>
|
| 71 |
|
| 72 |
-
> "
|
| 73 |
-
>
|
| 74 |
-
> the Church."
|
| 75 |
|
| 76 |
-
> "
|
| 77 |
-
>
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
## What it is good at / not good at
|
| 80 |
|
| 81 |
-
- **Good at:** producing grammatically
|
| 82 |
-
order, function words, and
|
| 83 |
-
|
| 84 |
-
- **Not good at:**
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
facts do not emerge in sampling. Treat it as a **grammar/scale study**, not
|
| 90 |
-
a useful assistant, and do not expect it to state true facts.
|
| 91 |
|
| 92 |
## Files
|
| 93 |
|
|
@@ -107,30 +105,18 @@ This is a custom architecture (not transformers-native). Load with the
|
|
| 107 |
```python
|
| 108 |
import sys, torch
|
| 109 |
sys.path.insert(0, "<path-to-this-repo>")
|
| 110 |
-
from train_compactlm5m import CompactLM
|
| 111 |
from tokenizers import Tokenizer
|
| 112 |
|
| 113 |
tok = Tokenizer.from_file("tokenizer.json")
|
| 114 |
-
|
| 115 |
|
| 116 |
from safetensors.torch import load_file
|
| 117 |
-
sd = load_file("model.safetensors")
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
```
|
| 125 |
-
|
| 126 |
-
## Reproducibility
|
| 127 |
-
|
| 128 |
-
Everything needed to reproduce is in this repo: the architecture class, the
|
| 129 |
-
training script, the eval script, the tokenizer, and the weights. The only
|
| 130 |
-
external dependency is the training corpus (fineweb-edu, streamed).
|
| 131 |
-
|
| 132 |
-
---
|
| 133 |
-
_Trained and published by @Compactbot for the small-language-model community.
|
| 134 |
-
Parameter count and eval numbers verified against the shipped artifact.
|
| 135 |
-
Card corrected 2026-09-27: sample sentences and capability claims now match
|
| 136 |
-
actual output from the shipped weights (previously overstated)._
|
|
|
|
| 10 |
- tiny-model
|
| 11 |
- slm
|
| 12 |
- small-language-model
|
|
|
|
| 13 |
- from-scratch
|
| 14 |
- llama-style
|
| 15 |
metrics:
|
|
|
|
| 51 |
- **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
|
| 52 |
dclm-baseline-1.0 second corpus failed to connect at build time on the
|
| 53 |
training host, so this run used a single corpus. Logged here honestly.
|
| 54 |
+
- **Budget:** ~100M tokens over a 30–50 min GPU window (RTX 5090).
|
| 55 |
- **Objective:** next-token cross-entropy.
|
| 56 |
|
| 57 |
## Results (measured, not asserted)
|
| 58 |
|
| 59 |
+
- **Validation loss:** 3.8775 (final checkpoint, step 20000)
|
| 60 |
+
- **Validation perplexity:** 48.30 (over held-out fineweb-edu text)
|
|
|
|
| 61 |
- **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
|
| 62 |
+
loop detector, max 3-gram fraction over the 40-word tail)
|
| 63 |
|
| 64 |
+
Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
|
| 65 |
+
`model.safetensors`**):
|
| 66 |
|
| 67 |
+
> "The cat sat on the heart, the body needs to do so. On the other hand, the
|
| 68 |
+
> heart is not able to control the heart's ability to stay quiet."
|
| 69 |
|
| 70 |
+
> "The sun rises in the air. The sun is still in the air and the sun is on the
|
| 71 |
+
> ground. The sun rises in the air and causes it to rise again."
|
|
|
|
| 72 |
|
| 73 |
+
> "Once upon a time when a patient has been exposed to a medical condition and
|
| 74 |
+
> is unable to diagnose a condition. The following are the following..."
|
| 75 |
+
|
| 76 |
+
These are representative of the model's actual output: grammatically
|
| 77 |
+
structured, on-topic at the sentence level, but semantically loose.
|
| 78 |
|
| 79 |
## What it is good at / not good at
|
| 80 |
|
| 81 |
+
- **Good at:** producing grammatically structured, on-topic English at the
|
| 82 |
+
sentence level. It knows common word order, function words, and some
|
| 83 |
+
world-fact associations.
|
| 84 |
+
- **Not good at:** sustained coherence over long passages, factual accuracy,
|
| 85 |
+
or general reasoning. At ~6M parameters and ~100M tokens the model captures
|
| 86 |
+
surface grammar and high-frequency associations but not stable semantics.
|
| 87 |
+
Longer generations drift. Treat it as a grammar/scale study, not a useful
|
| 88 |
+
assistant.
|
|
|
|
|
|
|
| 89 |
|
| 90 |
## Files
|
| 91 |
|
|
|
|
| 105 |
```python
|
| 106 |
import sys, torch
|
| 107 |
sys.path.insert(0, "<path-to-this-repo>")
|
| 108 |
+
from train_compactlm5m import CompactLM
|
| 109 |
from tokenizers import Tokenizer
|
| 110 |
|
| 111 |
tok = Tokenizer.from_file("tokenizer.json")
|
| 112 |
+
m = CompactLM(12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512).eval()
|
| 113 |
|
| 114 |
from safetensors.torch import load_file
|
| 115 |
+
sd = {k: v for k, v in load_file("model.safetensors").items()
|
| 116 |
+
if not k.startswith("head.weight")} # head.weight is tied to tok.weight
|
| 117 |
+
m.load_state_dict(sd, strict=False)
|
| 118 |
+
m.head.weight = m.tok.weight
|
| 119 |
+
|
| 120 |
+
ids = tok.encode("The cat sat on the").ids
|
| 121 |
+
# ... run m.forward on ids, sample, decode
|
| 122 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|