ChrisMcCormick commited on
Commit
7097693
·
verified ·
1 Parent(s): 1d67c5a

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # decoderstack-d12
2
+
3
+ Training-state checkpoints for the **DecoderStack d12** Colab notebook line
4
+ (`stacks/decoder-medium`, single-GPU 40GB-A100 configuration of the nanochat-style
5
+ d12: 286,261,730 params, 12 layers x 768, FA3 varlen, Muon + AdamW with per-step
6
+ schedule tables, bf16-live + uint16-mantissa fp32 masters).
7
+
8
+ ## Layout
9
+
10
+ ```
11
+ checkpoints/<run_name>/state_stepNNNNNN.pt
12
+ ```
13
+
14
+ Each `state_step*.pt` is the **entire** training state at that step, written by the
15
+ walkthrough notebook's `write_state` and consumed by its `load_state`:
16
+
17
+ - `params[name]` — per-Param `w` (live weights), `mantissa` (uint16 low bits of the
18
+ fp32 master, where applicable), `first_mntm`, `scnd_mntm`
19
+ - `step`, `t_step` — schedule position
20
+ - `rng` — torch CPU + CUDA generator states
21
+ - `batch` — the next unconsumed training micro-batch (`inputs`, `targets`,
22
+ `cu_seqlens`), so the walkthrough dissects exactly the batch training would have
23
+ seen next
24
+ - `config` — the full `StackConfig` (asserted on load)
25
+ - `code` — the notebook source that produced the state
26
+
27
+ Produced by the `DecoderStack d12 Walkthrough` notebook: Part 1 trains under the
28
+ real 1680-step schedule, stops at `cfg.walkthrough_step`, saves + pushes here; the
29
+ walkthrough part reloads and spells the last layer's forward/backward and one Muon
30
+ step out flat.
31
+
32
+ ## Current states (run `40GB-A100_d12_walkthrough`, 2026-08-31)
33
+
34
+ | file | where in the run | val bpb |
35
+ |---|---|---|
36
+ | `state_step000250.pt` | step 250 of 1680 (past all warmups, lr at peak) | 1.076657 |
37
+ | `state_step000015.pt` | step 15 (mid lr/scalar/lm-head warmup) | 1.905353 |
38
+
39
+ To walk through a different state, set `cfg.walkthrough_step` to its step number
40
+ and run the notebook's walkthrough part.