File size: 4,550 Bytes
fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c b0a40e0 99e3ef7 b0a40e0 1a279d1 b0a40e0 fc4ad9c 99e3ef7 fc4ad9c 8093b7b fc4ad9c b0a40e0 fc4ad9c b0a40e0 1a279d1 b0a40e0 8093b7b b0a40e0 fc4ad9c b0a40e0 fc4ad9c b0a40e0 fc4ad9c b0a40e0 fc4ad9c b0a40e0 fc4ad9c b0a40e0 fc4ad9c 8093b7b fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 1a279d1 fc4ad9c 99e3ef7 451f582 1a279d1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | ---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- SLM
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
- allenai/dclm-baseline
metrics:
- perplexity
---
# LDT-10M
A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.
**Requested by [DedeProGames](https://huggingface.co/DedeProGames) on the [model-requests board](https://huggingface.co/spaces/Compactbot/model-requests) (#12).**
## Architecture
| Parameter | Value |
|---|---|
| Params | **10,284,480** |
| Layers | 5 |
| d_model | 320 |
| Heads | 5 (MHA, GQA not used at this scale) |
| FFN dim | 896 (SwiGLU) |
| Vocab | 12,288 (gollem BPE) |
| Context | 512 |
| Embeddings | Tied (lm_head β tok.weight) |
| Norm | RMSNorm (eps 1e-5) |
| Attention | RoPE + causal SDPA |
| Dtype | float32 |
Standard LLaMA block: RMSNorm β MHA (RoPE) β residual β RMSNorm β SwiGLU FFN β residual.
## Training
| | v1 (first checkpoint) | v4 (current) |
|---|---|---|
| Steps | 4,000 | 16,000 (4,000 + 12,000 continued) |
| Batch size | 64 | 64 |
| LR | 3e-4 β 3e-5 (cosine) | 1e-4 β 1e-5 (cosine, fresh optimizer) |
| Data | FineWeb-Edu 30.1M + DCLM 21.4M = **~50.5M tok** | + FineWeb-Edu 150.6M + DCLM 107.3M = **~257.9M tok** |
| Cumulative tokens | ~50.5M | **~308M** |
| Tokens/param | ~4.9 | **~30.0** |
| Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) |
| Final val loss | 4.6020 (ppl 99.68) | **3.8943 (ppl 49.12)** |
### β οΈ Honest caveat: undertrained and still incoherent
DedeProGames requested **2.6B tokens**. This checkpoint is at **~308M tokens** β an **8.4Γ shortfall**. The GPU was occupied by other work for most of the training window.
At 30 tok/param the model has learned the **surface shape** of English much better than v1 β real words, parseable sentences, occasional token loops β but the prose is still **semantically incoherent** (word salad): grammatically plausible sentences that don't mean what they say. The val loss (3.89) is well below the 7.38 unigram floor, so it genuinely uses context; the improvement from 4.60 β 3.89 is real.
This is a **continued checkpoint**, not the final deliverable. More data is the fix; continued training toward the 2.6B budget is planned.
## Eval (8 prompts, temp 0.8, top-k 40, seed 1234)
| Metric | v1 | v4 |
|---|---|---|
| val loss | 4.6020 | **3.8943** |
| perplexity | 99.68 | **49.12** |
| Below unigram floor (7.38)? | Yes | Yes |
| Token loops? | No | Rare (1 hard loop in a 40-sample sweep, seeds 0-4) |
| Semantically coherent? | No | No (improved, still word salad) |
### Sample outputs (v4, real generation from the weights)
> "The cat sat on the same line of the moon as a dancer. The planet is the only planet that has been known to be a good deal of tear. The last thing about the world is a part of a real life."
> "Once upon a time, the person will be sent to each other, or if you are going to go through it. But a lot of people are going to have a good chance of a feeling. But that is, I do not know what they're going to do."
> "Water is made of hydrogen and oxygen, and it is a good deal of tear. The sun is a piece of light and is not a good deal of tear. The sun is not a good deal of tear."
> "To be or not to be is a question about the world. The last thing about the world is a part of a real life. It would be a great place, like it and in the earth."
These are real outputs from the v4 weights (not hand-picked for coherence). The text is grammatically structured β real words, parseable sentences, occasional token loops, no broken tokens β but semantically incoherent. That is the honest state of a 10M model at 30 tok/param.
## Usage
The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`).
To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_fixed.py`) contains the full architecture definition.
## What this is NOT
- Not a 2.6B-token model (that's the target; this is the ~308M checkpoint)
- Not a coherent-text model (it produces grammatically-structured word salad at this token budget)
- Not a general-purpose assistant (it's a raw LM, no instruction tuning)
- Not a replacement for anything larger β it's a research checkpoint in a from-scratch training run |