Commit History

Fix v4 token-count arithmetic: v4 data build is 257.9M tok (fw 150.6M + dclm 107.3M), cumulative ~308M, ~30 tok/param, 8.4x shortfall — not the previously stated ~393M/~444M/~43.1. Verified against models/ldt-10m-v4/train.log data-build lines.
d87c565
verified

Compactbot commited on

Update to v4 checkpoint: 444M tokens, val 3.8943 (ppl 49.12), still incoherent (#11)
b0a40e0

Compactbot commited on

v4 checkpoint: 444M tokens, val 3.8943 (ppl 49.12)
e68617a
verified

Compactbot commited on

Fix fabricated sample quote + overstatement: replace sample #1 (not in eval output) with real samples from eval_ldt10m_final.json; correct "coherent/fluency" to "grammatically structured but semantically incoherent" (#10)
451f582

Compactbot commited on

Add LDT-10M config (#9)
0ca7bb0

Compactbot commited on

Add LDT-10M weights (10,284,480 params, 47 tensors, tied embeddings) (#8)
44ce62e

Compactbot commited on

Add LDT-10M model card (#7)
1a279d1

Compactbot commited on

Add honest model card (architecture, training, undertraining status, real samples)
fc4ad9c
verified

Compactbot commited on

Tokenizer config (#5)
a7dc088

Compactbot commited on

Architecture config (#4)
efb909d

Compactbot commited on

Byte-level BPE tokenizer (vocab 12288) (#3)
2a1e042

Compactbot commited on

Self-contained LDT model definition (#2)
a198f75

Compactbot commited on

LDT-10M weights (47 tensors, 10,284,480 params, tied embeddings) (#1)
404b041

Compactbot commited on

initial commit
e116be6
verified

Compactbot commited on