iskhare commited on
Commit
dc1c07d
·
verified ·
1 Parent(s): 36aebcb

Initialize experiment namespace with training-budget provenance

Browse files
Files changed (2) hide show
  1. NCM-BUDGET-AUDIT.md +81 -0
  2. README.md +23 -0
NCM-BUDGET-AUDIT.md ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # NCM training-budget audit — 2026-09-16
2
+
3
+ ## Source revisions
4
+
5
+ - `HazyResearch/ncm`, `neurips-final-run`: `5f3e5641004de0f52e583b8841d79a9d538b3b63`.
6
+ - `HazyResearch/ncm`, `final`: `ab28dd9a9f67a729112feaf94147b26a5822e6a1`.
7
+ - HF `hazyresearch/ncm-neurips-owt`: `9570350baa885bd94abdaacff6c21675252d9c55`.
8
+
9
+ These are different configurations and historical runs, not a single definitive
10
+ NCM training budget. Matching the exact paper checkpoint requires its resolved
11
+ configuration and optimizer-step history, including world-size changes.
12
+
13
+ ## What the branches establish
14
+
15
+ The common `config/experiments/neurips/{vqvae-owt-gpt2-final,
16
+ ncp-owt-gpt2-final-small,ncp-owt-gpt2-final}.yaml` files specify **100 epochs**.
17
+ The generator configs actually have `context_length=256`, `prefix_len=128`,
18
+ and tokenizer context 128; some comments still say prefix 64 and are stale.
19
+ The models use 14 levels starting at 32, 50,304 codes/level, and 10% uniform
20
+ code-replacement corruption. This is not the current ctx256, prefix-free,
21
+ 12-level/16K-codebook reproduction.
22
+
23
+ `utils/arg_utils.py:parse_args` derives updates as:
24
+
25
+ ```
26
+ ceil(n_epochs * corpus_token_count /
27
+ (batch_per_gpu * model.context_length * accumulation * world_size))
28
+ ```
29
+
30
+ Therefore the 100-epoch YAML alone does not establish either executed steps or
31
+ an absolute token count without the actual dataset and runtime overrides.
32
+ The `final` branch also contains different context-specific recipes, e.g.
33
+ `config/owt-ncp-ctx128/ncp.yaml` specifies 50 epochs. Do not combine their
34
+ architecture, schedules, and checkpoint counts into one purported run.
35
+
36
+ ## What released checkpoint logs establish
37
+
38
+ HF paths `vqvae-owt/stdout.txt` and `ncp-owt-gpt2/stdout.txt` were inspected.
39
+
40
+ | Released artifact | Evidence | Interpretation |
41
+ |---|---|---|
42
+ | `vqvae-owt/checkpoint-iter-103152.pt` | Resolved log: 5 epochs, batch 512, accumulation 1, context 64; later resume warning explicitly identifies saved world size 4 | 103,152 × 512 × 4 × 64 = **13,520,338,944** raw positions for this checkpoint trajectory |
43
+ | `ncp-owt-gpt2/checkpoint-iter-80000.pt` | Batch 96, accumulation 1, context 128, text prefix 64; log contains both 8- and 4-GPU resumes | Step 80k is not 50 completed epochs. At fixed 4/8 GPUs, 80k updates correspond to **3.932/7.864B** raw positions, half of which are target positions. Exact cumulative accounting requires resolving the resume timeline |
44
+
45
+ The generator's resolved planned totals include 1,375,350 and 2,750,699 updates
46
+ with `n_epochs=50`. The released artifact is only step 80,000. The logs include
47
+ restarts/resumes; discarded computation is separate from the retained training
48
+ trajectory. Do not report the configured endpoint as completed training.
49
+
50
+ The released tokenizer count implies about 2.704B corpus tokens, consistent
51
+ with the current local uint16 binary (2,704,046,552 tokens; rounding to whole
52
+ updates explains the small difference). This is a size consistency check,
53
+ **not a dataset-content/hash identity claim**.
54
+
55
+ Repository evaluation records additionally reference checkpoints up to 140k
56
+ (`results/gen_ppl_ncp_owt_gpt2_full.json`) and a run named `300ep` through 105k
57
+ (`results/gen_ppl_ncp_rope_qknorm_corrupt_300ep.json`). Names/configured epochs
58
+ are not proof of 300 completed epochs or of their exact token budgets.
59
+
60
+ ## Current MsTok accounting
61
+
62
+ Batch 32/GPU × 12 accumulation × 8 GPUs × 256 = **786,432 positions/update**.
63
+ The current binary is 5,408,093,104 bytes of uint16 = 2,704,046,552 tokens.
64
+
65
+ | Updates per component | Repeated raw token positions | Nominal corpus passes |
66
+ |---|---:|---:|
67
+ | 10,000 | 7.864B | 2.91 |
68
+ | 17,192 | 13.520B | 5.00 |
69
+ | 25,000 | 19.661B | 7.27 |
70
+ | 34,384 | 27.041B | 10.00 |
71
+
72
+ The legacy 103,152-update tokenizer and our 17,192-update component therefore
73
+ have the **same raw-position budget**, despite a sixfold step-count difference.
74
+ At current batch/context, 131B positions require **166,576 updates**. At the
75
+ observed ~0.258 joint updates/second, that is roughly 7.5 days of training alone
76
+ on one eight-H100 node, excluding evaluations and overhead. More data is not
77
+ free and a multi-candidate sweep at that budget will not fit a ten-day deadline.
78
+
79
+ Before revising final-run lengths, identify which NCM checkpoint/result is the
80
+ comparison target. A short-budget tuning experiment remains useful, but is not
81
+ the same claim as matching a long-budget baseline.
README.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ tags:
5
+ - text-generation
6
+ - research
7
+ ---
8
+ # ICLR debug experiments
9
+
10
+ Public artifacts for the `iclr-debug` branch of HazyResearch/MsTok.
11
+ W&B project: `mstok/iclr-debug`.
12
+
13
+ This repository currently establishes the experiment namespace and contains a
14
+ training-budget audit. No new model weights or final benchmark claims are
15
+ implied. Future selected checkpoints and evaluations will be organized by
16
+ unique run ID and optimizer step.
17
+
18
+ Training budgets must report raw token positions per component, dataset size,
19
+ global batch, optimizer updates, and LR horizons. Generation reports must label
20
+ supplied-level-zero versus unconditional sampling and record the reference
21
+ model precision, entropy, scored-token counts, and per-seed results.
22
+
23
+ See `NCM-BUDGET-AUDIT.md` for historical budget provenance and limitations.