Fix over-stated quality claims: model is first-sentence-coherent then loops, not "fluent/on-topic"

#9
by Compactbot - opened
Files changed (1) hide show
  1. README.md +9 -7
README.md CHANGED
@@ -63,11 +63,13 @@ above is the honest held-out figure.
63
 
64
  ## What it is good at / not
65
 
66
- At 5M parameters this model produces fluent, grammatical, on-topic English,
67
- but it is a small model: factual recall is weak and greedy decoding drifts
68
- into repetition. Sampled decoding (top-p 0.9, temp 0.7) is noticeably more
69
- diverse. It is a demonstration of from-scratch small-model training, not a
70
- useful general assistant.
 
 
71
 
72
  ### Sample outputs (greedy, temp 0.0)
73
 
@@ -93,7 +95,7 @@ useful general assistant.
93
 
94
  ## Version
95
 
96
- This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) is the first version of this lineage that produces genuinely coherent, on-topic output, so it is published as the canonical model.
97
 
98
 
99
  - `model.safetensors` — weights (27 MB, F32). `head.weight` and `tok.weight`
@@ -106,4 +108,4 @@ This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d2
106
 
107
  Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
108
  script, tokenizer and training log are not bundled here; the architecture is
109
- fully specified in `config.json` and above.
 
63
 
64
  ## What it is good at / not
65
 
66
+ At 5M parameters this model produces grammatical first sentences but is far
67
+ from fluent: greedy decoding is coherent for roughly the first sentence and
68
+ then collapses into repetition loops (e.g. "the world's largest city in the
69
+ world is the world's largest city…"), and sampled decoding (top-p 0.9,
70
+ temp 0.7) is more varied but still drifts into incoherence within a few
71
+ sentences. Factual recall is weak. It is a demonstration of from-scratch
72
+ small-model training, not a useful general assistant.
73
 
74
  ### Sample outputs (greedy, temp 0.0)
75
 
 
95
 
96
  ## Version
97
 
98
+ This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.
99
 
100
 
101
  - `model.safetensors` — weights (27 MB, F32). `head.weight` and `tok.weight`
 
108
 
109
  Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
110
  script, tokenizer and training log are not bundled here; the architecture is
111
+ fully specified in `config.json` and above.