1gpu-llm Small V3 Base
Canonical decayed Small V3 release: native checkpoint step_41000, selected with MEDIUM confidence.
Role and lineage
acquisition โ acquisition release โ continual โ lower-LR continual โ pre-collapse step_39500 โ d2000 optimizer-preserving release decay โ final selection.
The release-decay branch restored the optimizer from step_39500, replaced the old scheduler with a fresh inverse-proportional decay-only scheduler, and ran 2,000 steps from 1.5e-5 to 1e-5. Dense checkpoints and the parent were evaluated before final selection.
Selection evidence
GPU/BF16 and CPU/FP32 scalar rankings were not fully concordant, and a bounded blind generation grid was ambiguous. Final adjudication was CPU/FP32 over frozen EN/IT/CODE held-out data with corpus weights EN 0.4464237878, IT 0.4464237878, CODE 0.1071524243, using 20,000 paired document bootstrap replicates.
40300conclusively beat40800on the weighted composite.40300and41000remained statistically overlapping;41000was not a direct-CI winner.- The preregistered minimax domain-regret tie-break selected
41000.
Weighted NLL: 40300=4.057829421505, 40800=4.080180181639, 41000=4.059070265073.
Maximum domain regret: 40300=0.053380883258, 40800=0.064487623132, 41000=0.028973489437.
Native artifact format
This repository publishes the original checkpoint (step_41000.pt) and an exact checkpoint-native safetensors export (step_41000.safetensors). The native artifacts are authoritative.
Standard Hugging Face Transformers model files are intentionally omitted. The current GPT-2 mapping drops the native head.bias; it has not been demonstrated numerically equivalent. Do not infer standard AutoModelForCausalLM compatibility from the tokenizer files.
Architecture and training identities
- GPT2-like PreLN causal LM: 12 layers, hidden size 768, 12 heads.
- 123,888,000 unique parameters; tied token embedding / LM head.
- Context length: 2,500 tokens; vocabulary: 48,000.
- Tokenizer SHA256:
8aef9589d7807ee613d338082ece6d2c71d7c684f114b640dd3a71f51224a08b. - Dataset: nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code at revision
333c4001551757aedb3454f30c8a6c08eaf23e12. - Mixture: English 6.666B, Italian 6.666B, code 1.600B training tokens.
- Checkpoint step:
41000; cursor:3936000; stored LR:1.0907107798582077e-05. - Cumulative exposure: 9.84B packed token-exposures; K: ~79.43.
- Original checkpoint SHA256:
1e58eea8a34764a673e682c534ad1b61a77cbd69eb5a55a22fd9e3416425c86f.
License
No license tag is asserted. The corpus release metadata does not establish a complete downstream license posture for the mixed upstream sources. Users are responsible for verifying applicable upstream terms before use or redistribution.
