1gpu-llm official model

1gpu-llm Small V3 Base

Canonical decayed Small V3 release: native checkpoint step_41000, selected with MEDIUM confidence.

Role and lineage

acquisition โ†’ acquisition release โ†’ continual โ†’ lower-LR continual โ†’ pre-collapse step_39500 โ†’ d2000 optimizer-preserving release decay โ†’ final selection.

The release-decay branch restored the optimizer from step_39500, replaced the old scheduler with a fresh inverse-proportional decay-only scheduler, and ran 2,000 steps from 1.5e-5 to 1e-5. Dense checkpoints and the parent were evaluated before final selection.

Selection evidence

GPU/BF16 and CPU/FP32 scalar rankings were not fully concordant, and a bounded blind generation grid was ambiguous. Final adjudication was CPU/FP32 over frozen EN/IT/CODE held-out data with corpus weights EN 0.4464237878, IT 0.4464237878, CODE 0.1071524243, using 20,000 paired document bootstrap replicates.

  • 40300 conclusively beat 40800 on the weighted composite.
  • 40300 and 41000 remained statistically overlapping; 41000 was not a direct-CI winner.
  • The preregistered minimax domain-regret tie-break selected 41000.

Weighted NLL: 40300=4.057829421505, 40800=4.080180181639, 41000=4.059070265073.

Maximum domain regret: 40300=0.053380883258, 40800=0.064487623132, 41000=0.028973489437.

Native artifact format

This repository publishes the original checkpoint (step_41000.pt) and an exact checkpoint-native safetensors export (step_41000.safetensors). The native artifacts are authoritative.

Standard Hugging Face Transformers model files are intentionally omitted. The current GPT-2 mapping drops the native head.bias; it has not been demonstrated numerically equivalent. Do not infer standard AutoModelForCausalLM compatibility from the tokenizer files.

Architecture and training identities

  • GPT2-like PreLN causal LM: 12 layers, hidden size 768, 12 heads.
  • 123,888,000 unique parameters; tied token embedding / LM head.
  • Context length: 2,500 tokens; vocabulary: 48,000.
  • Tokenizer SHA256: 8aef9589d7807ee613d338082ece6d2c71d7c684f114b640dd3a71f51224a08b.
  • Dataset: nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code at revision 333c4001551757aedb3454f30c8a6c08eaf23e12.
  • Mixture: English 6.666B, Italian 6.666B, code 1.600B training tokens.
  • Checkpoint step: 41000; cursor: 3936000; stored LR: 1.0907107798582077e-05.
  • Cumulative exposure: 9.84B packed token-exposures; K: ~79.43.
  • Original checkpoint SHA256: 1e58eea8a34764a673e682c534ad1b61a77cbd69eb5a55a22fd9e3416425c86f.

License

No license tag is asserted. The corpus release metadata does not establish a complete downstream license posture for the mixed upstream sources. Users are responsible for verifying applicable upstream terms before use or redistribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Dataset used to train nazdef/1gpu-llm-small-en-it-code-v3-base

Collection including nazdef/1gpu-llm-small-en-it-code-v3-base