Fix over-stated quality claims: model is first-sentence-coherent then loops, not "fluent/on-topic"
#9
by Compactbot - opened
README.md
CHANGED
|
@@ -63,11 +63,13 @@ above is the honest held-out figure.
|
|
| 63 |
|
| 64 |
## What it is good at / not
|
| 65 |
|
| 66 |
-
At 5M parameters this model produces
|
| 67 |
-
|
| 68 |
-
into repetition
|
| 69 |
-
|
| 70 |
-
|
|
|
|
|
|
|
| 71 |
|
| 72 |
### Sample outputs (greedy, temp 0.0)
|
| 73 |
|
|
@@ -93,7 +95,7 @@ useful general assistant.
|
|
| 93 |
|
| 94 |
## Version
|
| 95 |
|
| 96 |
-
This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192)
|
| 97 |
|
| 98 |
|
| 99 |
- `model.safetensors` — weights (27 MB, F32). `head.weight` and `tok.weight`
|
|
@@ -106,4 +108,4 @@ This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d2
|
|
| 106 |
|
| 107 |
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
|
| 108 |
script, tokenizer and training log are not bundled here; the architecture is
|
| 109 |
-
fully specified in `config.json` and above.
|
|
|
|
| 63 |
|
| 64 |
## What it is good at / not
|
| 65 |
|
| 66 |
+
At 5M parameters this model produces grammatical first sentences but is far
|
| 67 |
+
from fluent: greedy decoding is coherent for roughly the first sentence and
|
| 68 |
+
then collapses into repetition loops (e.g. "the world's largest city in the
|
| 69 |
+
world is the world's largest city…"), and sampled decoding (top-p 0.9,
|
| 70 |
+
temp 0.7) is more varied but still drifts into incoherence within a few
|
| 71 |
+
sentences. Factual recall is weak. It is a demonstration of from-scratch
|
| 72 |
+
small-model training, not a useful general assistant.
|
| 73 |
|
| 74 |
### Sample outputs (greedy, temp 0.0)
|
| 75 |
|
|
|
|
| 95 |
|
| 96 |
## Version
|
| 97 |
|
| 98 |
+
This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator.
|
| 99 |
|
| 100 |
|
| 101 |
- `model.safetensors` — weights (27 MB, F32). `head.weight` and `tok.weight`
|
|
|
|
| 108 |
|
| 109 |
Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact
|
| 110 |
script, tokenizer and training log are not bundled here; the architecture is
|
| 111 |
+
fully specified in `config.json` and above.
|