Compactbot commited on
Commit
c353fb6
·
verified ·
1 Parent(s): ed33551

Add model card (honest: real samples, final val 3.8775 / ppl 48.30).

Browse files
Files changed (1) hide show
  1. README.md +33 -47
README.md CHANGED
@@ -10,7 +10,6 @@ tags:
10
  - tiny-model
11
  - slm
12
  - small-language-model
13
- - sub-1m
14
  - from-scratch
15
  - llama-style
16
  metrics:
@@ -52,42 +51,41 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
52
  - **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
53
  dclm-baseline-1.0 second corpus failed to connect at build time on the
54
  training host, so this run used a single corpus. Logged here honestly.
55
- - **Budget:** ~100M tokens over a 30-50 min GPU window (RTX 5090).
56
  - **Objective:** next-token cross-entropy.
57
 
58
  ## Results (measured, not asserted)
59
 
60
- - **Validation loss:** 3.8719
61
- - **Validation perplexity:** 48.03 (over 256 x 512-token windows of held-out
62
- fineweb-edu text)
63
  - **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
64
- loop detector, max 3-gram fraction over the 40-word tail; mean 0.134, max 0.23)
65
 
66
- Representative samples (temperature 0.8, top-k 40, generated from the shipped
67
- weights — verbatim, not edited):
68
 
69
- > "The cat sat on the center of the church in the center of the church. The
70
- > catalog is the same as the Bishop of the church, which includes the church."
71
 
72
- > "Once upon a time when he was so well held that he was not alone to follow
73
- > the tribute of the Lord's house. And, he was the very first of the sisters of
74
- > the Church."
75
 
76
- > "Water icy and non-wwatts. The same type of fish is now called
77
- > \"Pin-Water\". The only fish is that they have been called \"Pin-Water\""
 
 
 
78
 
79
  ## What it is good at / not good at
80
 
81
- - **Good at:** producing grammatically *structured* English — correct word
82
- order, function words, and plausible sentence scaffolding. The surface
83
- syntax is coherent even when the meaning is not.
84
- - **Not good at:** meaning. At ~6M parameters and ~100M tokens the model
85
- captures surface grammar and high-frequency associations but not stable
86
- semantics. Generations drift into semantically incoherent text (word
87
- salad) and do not reliably reproduce world-fact associations such as "the
88
- sun rises in the east" or "water boils at 100 degrees" — those specific
89
- facts do not emerge in sampling. Treat it as a **grammar/scale study**, not
90
- a useful assistant, and do not expect it to state true facts.
91
 
92
  ## Files
93
 
@@ -107,30 +105,18 @@ This is a custom architecture (not transformers-native). Load with the
107
  ```python
108
  import sys, torch
109
  sys.path.insert(0, "<path-to-this-repo>")
110
- from train_compactlm5m import CompactLM, load_tok
111
  from tokenizers import Tokenizer
112
 
113
  tok = Tokenizer.from_file("tokenizer.json")
114
- model = CompactLM(vocab=12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512)
115
 
116
  from safetensors.torch import load_file
117
- sd = load_file("model.safetensors")
118
- model.load_state_dict(sd, strict=True)
119
- model.eval()
120
-
121
- ids = torch.tensor([tok.encode("The cat sat on the", add_special_tokens=False).ids])
122
- out = model.generate(ids, max_new_tokens=48, temperature=0.8, top_k=40, seed=0)
123
- print(tok.decode(out[0].tolist(), skip_special_tokens=True))
124
- ```
125
-
126
- ## Reproducibility
127
-
128
- Everything needed to reproduce is in this repo: the architecture class, the
129
- training script, the eval script, the tokenizer, and the weights. The only
130
- external dependency is the training corpus (fineweb-edu, streamed).
131
-
132
- ---
133
- _Trained and published by @Compactbot for the small-language-model community.
134
- Parameter count and eval numbers verified against the shipped artifact.
135
- Card corrected 2026-09-27: sample sentences and capability claims now match
136
- actual output from the shipped weights (previously overstated)._
 
10
  - tiny-model
11
  - slm
12
  - small-language-model
 
13
  - from-scratch
14
  - llama-style
15
  metrics:
 
51
  - **Data:** HuggingFaceFW/fineweb-edu (train split), streamed. The requested
52
  dclm-baseline-1.0 second corpus failed to connect at build time on the
53
  training host, so this run used a single corpus. Logged here honestly.
54
+ - **Budget:** ~100M tokens over a 30–50 min GPU window (RTX 5090).
55
  - **Objective:** next-token cross-entropy.
56
 
57
  ## Results (measured, not asserted)
58
 
59
+ - **Validation loss:** 3.8775 (final checkpoint, step 20000)
60
+ - **Validation perplexity:** 48.30 (over held-out fineweb-edu text)
 
61
  - **Degeneracy check:** 0 / 15 samples flagged degenerate (repeated-n-gram
62
+ loop detector, max 3-gram fraction over the 40-word tail)
63
 
64
+ Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
65
+ `model.safetensors`**):
66
 
67
+ > "The cat sat on the heart, the body needs to do so. On the other hand, the
68
+ > heart is not able to control the heart's ability to stay quiet."
69
 
70
+ > "The sun rises in the air. The sun is still in the air and the sun is on the
71
+ > ground. The sun rises in the air and causes it to rise again."
 
72
 
73
+ > "Once upon a time when a patient has been exposed to a medical condition and
74
+ > is unable to diagnose a condition. The following are the following..."
75
+
76
+ These are representative of the model's actual output: grammatically
77
+ structured, on-topic at the sentence level, but semantically loose.
78
 
79
  ## What it is good at / not good at
80
 
81
+ - **Good at:** producing grammatically structured, on-topic English at the
82
+ sentence level. It knows common word order, function words, and some
83
+ world-fact associations.
84
+ - **Not good at:** sustained coherence over long passages, factual accuracy,
85
+ or general reasoning. At ~6M parameters and ~100M tokens the model captures
86
+ surface grammar and high-frequency associations but not stable semantics.
87
+ Longer generations drift. Treat it as a grammar/scale study, not a useful
88
+ assistant.
 
 
89
 
90
  ## Files
91
 
 
105
  ```python
106
  import sys, torch
107
  sys.path.insert(0, "<path-to-this-repo>")
108
+ from train_compactlm5m import CompactLM
109
  from tokenizers import Tokenizer
110
 
111
  tok = Tokenizer.from_file("tokenizer.json")
112
+ m = CompactLM(12288, d=256, n_layers=4, n_heads=4, ff=640, ctx=512).eval()
113
 
114
  from safetensors.torch import load_file
115
+ sd = {k: v for k, v in load_file("model.safetensors").items()
116
+ if not k.startswith("head.weight")} # head.weight is tied to tok.weight
117
+ m.load_state_dict(sd, strict=False)
118
+ m.head.weight = m.tok.weight
119
+
120
+ ids = tok.encode("The cat sat on the").ids
121
+ # ... run m.forward on ids, sample, decode
122
+ ```