ldt-10m / README.md
Compactbot's picture
Fix over-strong 'no token loops' claim: 40-sample sweep (seeds 0-4) found 1 hard loop (seed 4, loop_frac 1.0) + several elevated-loop samples. Card now says 'occasional/rare token loops' instead of 'no token loops'. All other numbers (val 3.8943, ppl 49.12, ~308M tok) re-verified and unchanged.
8093b7b verified
|
Raw History Blame
4.55 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - SLM
  - small-language-model
  - from-scratch
  - llama
datasets:
  - HuggingFaceFW/fineweb-edu
  - allenai/dclm-baseline
metrics:
  - perplexity

LDT-10M

A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.

Requested by DedeProGames on the model-requests board (#12).

Architecture

Parameter Value
Params 10,284,480
Layers 5
d_model 320
Heads 5 (MHA, GQA not used at this scale)
FFN dim 896 (SwiGLU)
Vocab 12,288 (gollem BPE)
Context 512
Embeddings Tied (lm_head β†’ tok.weight)
Norm RMSNorm (eps 1e-5)
Attention RoPE + causal SDPA
Dtype float32

Standard LLaMA block: RMSNorm β†’ MHA (RoPE) β†’ residual β†’ RMSNorm β†’ SwiGLU FFN β†’ residual.

Training

v1 (first checkpoint) v4 (current)
Steps 4,000 16,000 (4,000 + 12,000 continued)
Batch size 64 64
LR 3e-4 β†’ 3e-5 (cosine) 1e-4 β†’ 1e-5 (cosine, fresh optimizer)
Data FineWeb-Edu 30.1M + DCLM 21.4M = ~50.5M tok + FineWeb-Edu 150.6M + DCLM 107.3M = ~257.9M tok
Cumulative tokens ~50.5M ~308M
Tokens/param ~4.9 ~30.0
Hardware RTX 5090 (32 GB) RTX 5090 (32 GB)
Final val loss 4.6020 (ppl 99.68) 3.8943 (ppl 49.12)

⚠️ Honest caveat: undertrained and still incoherent

DedeProGames requested 2.6B tokens. This checkpoint is at ~308M tokens β€” an 8.4Γ— shortfall. The GPU was occupied by other work for most of the training window.

At 30 tok/param the model has learned the surface shape of English much better than v1 β€” real words, parseable sentences, occasional token loops β€” but the prose is still semantically incoherent (word salad): grammatically plausible sentences that don't mean what they say. The val loss (3.89) is well below the 7.38 unigram floor, so it genuinely uses context; the improvement from 4.60 β†’ 3.89 is real.

This is a continued checkpoint, not the final deliverable. More data is the fix; continued training toward the 2.6B budget is planned.

Eval (8 prompts, temp 0.8, top-k 40, seed 1234)

Metric v1 v4
val loss 4.6020 3.8943
perplexity 99.68 49.12
Below unigram floor (7.38)? Yes Yes
Token loops? No Rare (1 hard loop in a 40-sample sweep, seeds 0-4)
Semantically coherent? No No (improved, still word salad)

Sample outputs (v4, real generation from the weights)

"The cat sat on the same line of the moon as a dancer. The planet is the only planet that has been known to be a good deal of tear. The last thing about the world is a part of a real life."

"Once upon a time, the person will be sent to each other, or if you are going to go through it. But a lot of people are going to have a good chance of a feeling. But that is, I do not know what they're going to do."

"Water is made of hydrogen and oxygen, and it is a good deal of tear. The sun is a piece of light and is not a good deal of tear. The sun is not a good deal of tear."

"To be or not to be is a question about the world. The last thing about the world is a part of a real life. It would be a great place, like it and in the earth."

These are real outputs from the v4 weights (not hand-picked for coherence). The text is grammatically structured β€” real words, parseable sentences, occasional token loops, no broken tokens β€” but semantically incoherent. That is the honest state of a 10M model at 30 tok/param.

Usage

The model uses a custom LDT architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares tok.weight).

To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (train_ldt10m_fixed.py) contains the full architecture definition.

What this is NOT

  • Not a 2.6B-token model (that's the target; this is the ~308M checkpoint)
  • Not a coherent-text model (it produces grammatically-structured word salad at this token budget)
  • Not a general-purpose assistant (it's a raw LM, no instruction tuning)
  • Not a replacement for anything larger β€” it's a research checkpoint in a from-scratch training run