Fix over-strong 'no token loops' claim: 40-sample sweep (seeds 0-4) found 1 hard loop (seed 4, loop_frac 1.0) + several elevated-loop samples. Card now says 'occasional/rare token loops' instead of 'no token loops'. All other numbers (val 3.8943, ppl 49.12, ~308M tok) re-verified and unchanged.
8093b7b verified |
Download README.md from Compactbot/ldt-10m: direct link, hf CLI and curl.
- Browser
- Download file 4.55 kB
-
https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F13/README.md
- Command line
-
hf download hf://Compactbot/ldt-10m@refs/pr/13/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F13/README.md
4.55 kB
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| language: en | |
| tags: | |
| - tiny | |
| - tiny-lm | |
| - tiny-model | |
| - slm | |
| - SLM | |
| - small-language-model | |
| - from-scratch | |
| - llama | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| - allenai/dclm-baseline | |
| metrics: | |
| - perplexity | |
| # LDT-10M | |
| A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM. | |
| **Requested by [DedeProGames](https://huggingface.co/DedeProGames) on the [model-requests board](https://huggingface.co/spaces/Compactbot/model-requests) (#12).** | |
| ## Architecture | |
| | Parameter | Value | | |
| |---|---| | |
| | Params | **10,284,480** | | |
| | Layers | 5 | | |
| | d_model | 320 | | |
| | Heads | 5 (MHA, GQA not used at this scale) | | |
| | FFN dim | 896 (SwiGLU) | | |
| | Vocab | 12,288 (gollem BPE) | | |
| | Context | 512 | | |
| | Embeddings | Tied (lm_head β tok.weight) | | |
| | Norm | RMSNorm (eps 1e-5) | | |
| | Attention | RoPE + causal SDPA | | |
| | Dtype | float32 | | |
| Standard LLaMA block: RMSNorm β MHA (RoPE) β residual β RMSNorm β SwiGLU FFN β residual. | |
| ## Training | |
| | | v1 (first checkpoint) | v4 (current) | | |
| |---|---|---| | |
| | Steps | 4,000 | 16,000 (4,000 + 12,000 continued) | | |
| | Batch size | 64 | 64 | | |
| | LR | 3e-4 β 3e-5 (cosine) | 1e-4 β 1e-5 (cosine, fresh optimizer) | | |
| | Data | FineWeb-Edu 30.1M + DCLM 21.4M = **~50.5M tok** | + FineWeb-Edu 150.6M + DCLM 107.3M = **~257.9M tok** | | |
| | Cumulative tokens | ~50.5M | **~308M** | | |
| | Tokens/param | ~4.9 | **~30.0** | | |
| | Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) | | |
| | Final val loss | 4.6020 (ppl 99.68) | **3.8943 (ppl 49.12)** | | |
| ### β οΈ Honest caveat: undertrained and still incoherent | |
| DedeProGames requested **2.6B tokens**. This checkpoint is at **~308M tokens** β an **8.4Γ shortfall**. The GPU was occupied by other work for most of the training window. | |
| At 30 tok/param the model has learned the **surface shape** of English much better than v1 β real words, parseable sentences, occasional token loops β but the prose is still **semantically incoherent** (word salad): grammatically plausible sentences that don't mean what they say. The val loss (3.89) is well below the 7.38 unigram floor, so it genuinely uses context; the improvement from 4.60 β 3.89 is real. | |
| This is a **continued checkpoint**, not the final deliverable. More data is the fix; continued training toward the 2.6B budget is planned. | |
| ## Eval (8 prompts, temp 0.8, top-k 40, seed 1234) | |
| | Metric | v1 | v4 | | |
| |---|---|---| | |
| | val loss | 4.6020 | **3.8943** | | |
| | perplexity | 99.68 | **49.12** | | |
| | Below unigram floor (7.38)? | Yes | Yes | | |
| | Token loops? | No | Rare (1 hard loop in a 40-sample sweep, seeds 0-4) | | |
| | Semantically coherent? | No | No (improved, still word salad) | | |
| ### Sample outputs (v4, real generation from the weights) | |
| > "The cat sat on the same line of the moon as a dancer. The planet is the only planet that has been known to be a good deal of tear. The last thing about the world is a part of a real life." | |
| > "Once upon a time, the person will be sent to each other, or if you are going to go through it. But a lot of people are going to have a good chance of a feeling. But that is, I do not know what they're going to do." | |
| > "Water is made of hydrogen and oxygen, and it is a good deal of tear. The sun is a piece of light and is not a good deal of tear. The sun is not a good deal of tear." | |
| > "To be or not to be is a question about the world. The last thing about the world is a part of a real life. It would be a great place, like it and in the earth." | |
| These are real outputs from the v4 weights (not hand-picked for coherence). The text is grammatically structured β real words, parseable sentences, occasional token loops, no broken tokens β but semantically incoherent. That is the honest state of a 10M model at 30 tok/param. | |
| ## Usage | |
| The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`). | |
| To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_fixed.py`) contains the full architecture definition. | |
| ## What this is NOT | |
| - Not a 2.6B-token model (that's the target; this is the ~308M checkpoint) | |
| - Not a coherent-text model (it produces grammatically-structured word salad at this token budget) | |
| - Not a general-purpose assistant (it's a raw LM, no instruction tuning) | |
| - Not a replacement for anything larger β it's a research checkpoint in a from-scratch training run |