1gpu-llm Small EN/IT/CODE β Acquisition-Release Base (step_11900)
This repository is the canonical Small base release for the 1gpu-llm EN/IT/CODE 15B, tokenizer-48k, ctx2500 lineage: the benchmark-selected winner of the acquisition-release decay branch.
1gpu-llm is a family of language models trained from scratch on a single consumer GPU.
Reference training hardware: NVIDIA GeForce RTX 4060 Ti 16GB, single GPU.
Checkpoint identity:
- family
1gpu-llm, tiersmall, role: decayed acquisition-release base - EN / IT / code, trained from scratch
- GPT-2-style decoder, pre-layernorm, tied embeddings, learned absolute positions
- vocab
48000, context2500, dim768, layers12, heads12 - parameters
160,752,000(~160.8M) - global step
11900, stored LR2.389e-05 - SHA256:
d5ce343cb05b34c25e33502cd39fddf85de00bbd428d1e893127d645dbacc1ab
This is a base pretrained model, not instruction tuned.
Training Lineage
acquisition (WSD, peak LR 2e-4)
-> best pre-collapse step_11000 (val_loss_mixed 3.3720)
-> optimizer-preserving / scheduler-reset decay-only branch
(1100 steps, 2e-4 -> 2e-5, inverse-proportional, ckpt every 100)
-> dense checkpoints 11100...12100
-> repo-native full benchmark (12 candidates)
-> winner step_11900 (900 decay steps in, NOT the 12100 endpoint)
Parent acquisition checkpoint: step_11000.pt, SHA256 19c7bf81β¦fc4b,
published separately as nazdef/1gpu-llm-small-en-it-code-acquisition-step11000
(revision 26e6479763e4612c7babbd571cd40b03d17b7527).
Selection Evidence
Full cohort (cuda/bf16/bs4/seed 1337, suite 20260919_pretrain_minimal_48k_en_it_code):
step_11900: mixed 4.9239, en 4.6377, it 3.9426, ppl 137.5- runner-up
step_11700: mixed 4.9359 (+0.012) - vs source
step_11000: mixed 5.1550 β β4.5%
Shortlist CPU/FP32 confirmation (11700 / 11900 / 12100, same suite+seed):
step_11900: mixed 4.9193, en 4.6333, it 3.9477, ppl 136.9, loop 0.375, lc 0.950/0.775step_11700: mixed 4.9375 (+0.018) βstep_12100: mixed 4.9522 (+0.033)
11900 won mixed loss under both contracts and holds the best EN in both. Small honest trade-off: IT is ~0.03 better at 11700/12100. No decoding grid was needed.
Relation to the Non-Decayed Artifact
Both are intentionally retained:
step_11000(acquisition repo): non-decayed, plastic, SFT-source candidatestep_11900(this repo): decayed, cooled, ready-to-use base and future Continual parent
Approximate Exposure
Project accounting (approximate exposure, not unique coverage):
240,000sequence tokens / optimizer step- ~2.856B sequence tokens by step 11900
The public dataset is document-level; training used its deterministic 48k/ctx2500 packed derivative (manifest 1e8b33f3β¦, shuffle pcg64/1337). Frozen dataset revision: 333c4001551757aedb3454f30c8a6c08eaf23e12.
Files Included
step_11900.ptβ original exact checkpoint (model + optimizer + scheduler + RNG + cursor 1142400)step_11900.safetensorsβ checkpoint-native weights, bit-identical (max delta 0.0)step_11900.safetensors.jsonβ export manifest / provenance sidecar- exact tokenizer bundle (SHA256
8aef9589β¦24a08b) source_decay_launch_config.yamlβ decay config that produced this branchrelease_selection.jsonβ GPU/BF16 + CPU/FP32 selection numbersrelease_manifest.json,SHA256SUMS
No standard-Transformers model.safetensors / config.json yet (see limitation).
Usage (repo-native)
import torch
from safetensors.torch import load_file
state = load_file("step_11900.safetensors") # exact model tensors
Native architecture: repo GPT2DecoderLM (pre-LN, fused QKV); load with project code at github.com/nazarenodefrancesc/nanochat-llm-training.
Limitations
- base pretrained model, not instruction tuned
- native format is authoritative; standard Transformers conversion omitted because the GPT-2 mapping drops
head.bias(max abs 1.327, mean abs 0.058) - benchmark is limited and does not prove factual reliability
- EN/IT/code balance is not uniform (IT ~0.03 better at neighboring checkpoints)
- downstream deployment requires separate evaluation
License / Data Provenance
No blanket license claim: the upstream licensing of the 15B EN/IT/CODE corpus mixture could not be established from the dataset release metadata (no license field on the dataset card). Same honest caveat as the acquisition repo. Users must verify upstream terms against the frozen revision above.
Summary
If you want the ready-to-use decayed Small base of this lineage β benchmark-selected under two contracts, cooled gradients, exact native weights β this is the one.