File size: 5,584 Bytes
fc4ad9c
1a279d1
fc4ad9c
 
 
 
1a279d1
fc4ad9c
 
1a279d1
fc4ad9c
 
1a279d1
 
 
 
fc4ad9c
 
 
 
 
 
1a279d1
fc4ad9c
1a279d1
fc4ad9c
1a279d1
fc4ad9c
1a279d1
fc4ad9c
1a279d1
 
fc4ad9c
1a279d1
 
 
 
 
 
 
 
 
 
fc4ad9c
 
 
c0527a5
 
 
 
 
 
 
 
 
 
1a279d1
c0527a5
fc4ad9c
c0527a5
fc4ad9c
c0527a5
fc4ad9c
c0527a5
fc4ad9c
e9f874b
fc4ad9c
c0527a5
 
 
 
 
 
 
 
fc4ad9c
c0527a5
fc4ad9c
c0527a5
fc4ad9c
c0527a5
fc4ad9c
c0527a5
 
 
fc4ad9c
1a279d1
fc4ad9c
1a279d1
fc4ad9c
c0527a5
fc4ad9c
e9f874b
 
 
 
 
 
 
 
 
 
 
 
1a279d1
fc4ad9c
c0527a5
1a279d1
c0527a5
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - SLM
  - small-language-model
  - from-scratch
  - llama
datasets:
  - HuggingFaceFW/fineweb-edu
  - allenai/dclm-baseline
metrics:
  - perplexity
---

# LDT-10M

A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.

**Requested by [DedeProGames](https://huggingface.co/DedeProGames) on the [model-requests board](https://huggingface.co/spaces/Compactbot/model-requests) (#12).**

## Architecture

| Parameter | Value |
|---|---|
| Params | **10,284,480** |
| Layers | 5 |
| d_model | 320 |
| Heads | 5 (MHA, GQA not used at this scale) |
| FFN dim | 896 (SwiGLU) |
| Vocab | 12,288 (gollem BPE) |
| Context | 512 |
| Embeddings | Tied (lm_head β†’ tok.weight) |
| Norm | RMSNorm (eps 1e-5) |
| Attention | RoPE + causal SDPA |
| Dtype | float32 |

Standard LLaMA block: RMSNorm β†’ MHA (RoPE) β†’ residual β†’ RMSNorm β†’ SwiGLU FFN β†’ residual.

## Training

| | v1 (first ckpt) | v4 | v7 | **v8 (current)** |
|---|---|---|---|---|
| Steps | 4,000 | 16,000 | 30,000 | **79,375** |
| Batch size | 64 | 64 | 64 | 64 |
| Seq length | 512 | 512 | 512 | 512 |
| Cumulative tokens | ~131M | ~308M | ~983M | **2.60B** |
| LR | 3e-4 β†’ 3e-5 (cosine) | 1e-4 β†’ 1e-5 (cosine, fresh optimizer) | 1e-4 β†’ 1e-5 (cosine, fresh optimizer) | 1e-4 β†’ 1e-5 (cosine, fresh optimizer) |
| Data | FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM (continued) | + FineWeb-Edu + DCLM (continued) |
| Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) | RTX 5090 (32 GB) | RTX 5090 (32 GB) |
| Final val loss | 4.6020 (ppl 99.68) | 3.8943 (ppl 49.12) | 3.6892 (ppl 40.01) | **3.63 (ppl 37.8)** |

v8 is the final checkpoint of the from-scratch run: 79,375 steps Γ— 64 Γ— 512 = **2,600,960,000 tokens (2.60B)**, which **hits DedeProGames' 2.6B-token target**. The val loss is read from the final checkpoint's recorded `val_loss` (3.63); the earlier columns' token counts are step-derived (steps Γ— 64 Γ— 512).

### ⚠️ Honest caveat: token target hit, but still incoherent and loop-prone

v8 reaches the requested 2.6B-token budget. The val loss (3.63) is well below the 7.38 unigram floor, so the model genuinely uses context; the improvement 4.60 β†’ 3.89 β†’ 3.69 β†’ 3.63 is real but **marginal in the last stage** (3.69 β†’ 3.63).

However, the 40-sample generation sweep below shows that **more tokens did not buy coherence β€” it bought more token loops.** The mean 4-gram loop fraction went **up** from v7 (0.167) to v8 (0.219), and the degenerate-sample rate (loop > 0.30) went from 0/40 to **14/40**. At this scale the model has learned the surface shape of English β€” real words, parseable sentence frames β€” but the prose is still **semantically incoherent** (word salad) and increasingly prone to hard repetition loops. This is the honest state of a 10M model at 2.6B tokens: the token budget is spent, and the quality ceiling of a 10M-param LM is what it is.

## Eval (40 samples: 8 prompts Γ— 5 seeds, temp 0.8, top-k 40, 160 new tokens)

| Metric | v1 | v4 | v7 | **v8** |
|---|---|---|---|---|
| val loss | 4.6020 | 3.8943 | 3.6892 | **3.63** |
| perplexity | 99.68 | 49.12 | 40.01 | **37.8** |
| Below unigram floor (7.38)? | Yes | Yes | Yes | Yes |
| mean 4-gram loop fraction | n/a | ~0.11 | 0.167 | **0.219** |
| Degenerate samples (loop > 0.30) | n/a | 0/40 | 0/40 | **14/40** |
| Semantically coherent? | No | No | No | No (word salad, more loops) |

### Sample outputs (v8, real generation from the published weights β€” not hand-picked)

> "Once upon a time, the time of the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord" (greedy)

> "She walked down the street. When she said she said she said she said she said she said she said she said she said she said she wa…" (temp 0.8, seed 0, loop frac 0.595)

> "A long time ago in a galaxy far far away from an galaxy that is almost no point from the distant distant galaxy called a…" (temp 0.8, seed 2, loop frac 0.587)

These are real outputs from the v8 weights. The text is grammatically structured β€” real words, parseable sentence frames β€” but semantically incoherent and loop-prone. That is the honest state of a 10M model at this scale.

## Usage

The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`).

To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_v8.py`) contains the full architecture definition.

```python
from model import LDT
from tokenizers import Tokenizer

model = LDT.from_pretrained("model.safetensors")  # re-binds the tied lm_head
tok = Tokenizer.from_file("tokenizer.json")

ids = tok.encode("The sun is", add_special_tokens=False).ids
out = model.generate(ids.unsqueeze(0), max_new_tokens=60, temperature=0.8, top_k=40, seed=0)
print(tok.decode(out[0].tolist(), skip_special_tokens=True))
```

## What this is NOT

- Not a coherent-text model (it produces grammatically-structured word salad with frequent token loops at this scale)
- Not a general-purpose assistant (it's a raw LM, no instruction tuning)
- Not a replacement for anything larger β€” it is the final from-scratch checkpoint of a 10M-param run that hit its 2.6B-token budget