Download README.md from Compactbot/ldt-10m: direct link, hf CLI and curl.
- Browser
- Download file 4.55 kB
-
https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F13/README.md
- Command line
-
hf download hf://Compactbot/ldt-10m@refs/pr/13/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F13/README.md
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- SLM
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
- allenai/dclm-baseline
metrics:
- perplexity
LDT-10M
A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.
Requested by DedeProGames on the model-requests board (#12).
Architecture
| Parameter | Value |
|---|---|
| Params | 10,284,480 |
| Layers | 5 |
| d_model | 320 |
| Heads | 5 (MHA, GQA not used at this scale) |
| FFN dim | 896 (SwiGLU) |
| Vocab | 12,288 (gollem BPE) |
| Context | 512 |
| Embeddings | Tied (lm_head β tok.weight) |
| Norm | RMSNorm (eps 1e-5) |
| Attention | RoPE + causal SDPA |
| Dtype | float32 |
Standard LLaMA block: RMSNorm β MHA (RoPE) β residual β RMSNorm β SwiGLU FFN β residual.
Training
| v1 (first checkpoint) | v4 (current) | |
|---|---|---|
| Steps | 4,000 | 16,000 (4,000 + 12,000 continued) |
| Batch size | 64 | 64 |
| LR | 3e-4 β 3e-5 (cosine) | 1e-4 β 1e-5 (cosine, fresh optimizer) |
| Data | FineWeb-Edu 30.1M + DCLM 21.4M = ~50.5M tok | + FineWeb-Edu 150.6M + DCLM 107.3M = ~257.9M tok |
| Cumulative tokens | ~50.5M | ~308M |
| Tokens/param | ~4.9 | ~30.0 |
| Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) |
| Final val loss | 4.6020 (ppl 99.68) | 3.8943 (ppl 49.12) |
β οΈ Honest caveat: undertrained and still incoherent
DedeProGames requested 2.6B tokens. This checkpoint is at ~308M tokens β an 8.4Γ shortfall. The GPU was occupied by other work for most of the training window.
At 30 tok/param the model has learned the surface shape of English much better than v1 β real words, parseable sentences, occasional token loops β but the prose is still semantically incoherent (word salad): grammatically plausible sentences that don't mean what they say. The val loss (3.89) is well below the 7.38 unigram floor, so it genuinely uses context; the improvement from 4.60 β 3.89 is real.
This is a continued checkpoint, not the final deliverable. More data is the fix; continued training toward the 2.6B budget is planned.
Eval (8 prompts, temp 0.8, top-k 40, seed 1234)
| Metric | v1 | v4 |
|---|---|---|
| val loss | 4.6020 | 3.8943 |
| perplexity | 99.68 | 49.12 |
| Below unigram floor (7.38)? | Yes | Yes |
| Token loops? | No | Rare (1 hard loop in a 40-sample sweep, seeds 0-4) |
| Semantically coherent? | No | No (improved, still word salad) |
Sample outputs (v4, real generation from the weights)
"The cat sat on the same line of the moon as a dancer. The planet is the only planet that has been known to be a good deal of tear. The last thing about the world is a part of a real life."
"Once upon a time, the person will be sent to each other, or if you are going to go through it. But a lot of people are going to have a good chance of a feeling. But that is, I do not know what they're going to do."
"Water is made of hydrogen and oxygen, and it is a good deal of tear. The sun is a piece of light and is not a good deal of tear. The sun is not a good deal of tear."
"To be or not to be is a question about the world. The last thing about the world is a part of a real life. It would be a great place, like it and in the earth."
These are real outputs from the v4 weights (not hand-picked for coherence). The text is grammatically structured β real words, parseable sentences, occasional token loops, no broken tokens β but semantically incoherent. That is the honest state of a 10M model at 30 tok/param.
Usage
The model uses a custom LDT architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares tok.weight).
To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (train_ldt10m_fixed.py) contains the full architecture definition.
What this is NOT
- Not a 2.6B-token model (that's the target; this is the ~308M checkpoint)
- Not a coherent-text model (it produces grammatically-structured word salad at this token budget)
- Not a general-purpose assistant (it's a raw LM, no instruction tuning)
- Not a replacement for anything larger β it's a research checkpoint in a from-scratch training run