Initialize experiment namespace with training-budget provenance
Browse files- NCM-BUDGET-AUDIT.md +81 -0
- README.md +23 -0
NCM-BUDGET-AUDIT.md
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# NCM training-budget audit — 2026-09-16
|
| 2 |
+
|
| 3 |
+
## Source revisions
|
| 4 |
+
|
| 5 |
+
- `HazyResearch/ncm`, `neurips-final-run`: `5f3e5641004de0f52e583b8841d79a9d538b3b63`.
|
| 6 |
+
- `HazyResearch/ncm`, `final`: `ab28dd9a9f67a729112feaf94147b26a5822e6a1`.
|
| 7 |
+
- HF `hazyresearch/ncm-neurips-owt`: `9570350baa885bd94abdaacff6c21675252d9c55`.
|
| 8 |
+
|
| 9 |
+
These are different configurations and historical runs, not a single definitive
|
| 10 |
+
NCM training budget. Matching the exact paper checkpoint requires its resolved
|
| 11 |
+
configuration and optimizer-step history, including world-size changes.
|
| 12 |
+
|
| 13 |
+
## What the branches establish
|
| 14 |
+
|
| 15 |
+
The common `config/experiments/neurips/{vqvae-owt-gpt2-final,
|
| 16 |
+
ncp-owt-gpt2-final-small,ncp-owt-gpt2-final}.yaml` files specify **100 epochs**.
|
| 17 |
+
The generator configs actually have `context_length=256`, `prefix_len=128`,
|
| 18 |
+
and tokenizer context 128; some comments still say prefix 64 and are stale.
|
| 19 |
+
The models use 14 levels starting at 32, 50,304 codes/level, and 10% uniform
|
| 20 |
+
code-replacement corruption. This is not the current ctx256, prefix-free,
|
| 21 |
+
12-level/16K-codebook reproduction.
|
| 22 |
+
|
| 23 |
+
`utils/arg_utils.py:parse_args` derives updates as:
|
| 24 |
+
|
| 25 |
+
```
|
| 26 |
+
ceil(n_epochs * corpus_token_count /
|
| 27 |
+
(batch_per_gpu * model.context_length * accumulation * world_size))
|
| 28 |
+
```
|
| 29 |
+
|
| 30 |
+
Therefore the 100-epoch YAML alone does not establish either executed steps or
|
| 31 |
+
an absolute token count without the actual dataset and runtime overrides.
|
| 32 |
+
The `final` branch also contains different context-specific recipes, e.g.
|
| 33 |
+
`config/owt-ncp-ctx128/ncp.yaml` specifies 50 epochs. Do not combine their
|
| 34 |
+
architecture, schedules, and checkpoint counts into one purported run.
|
| 35 |
+
|
| 36 |
+
## What released checkpoint logs establish
|
| 37 |
+
|
| 38 |
+
HF paths `vqvae-owt/stdout.txt` and `ncp-owt-gpt2/stdout.txt` were inspected.
|
| 39 |
+
|
| 40 |
+
| Released artifact | Evidence | Interpretation |
|
| 41 |
+
|---|---|---|
|
| 42 |
+
| `vqvae-owt/checkpoint-iter-103152.pt` | Resolved log: 5 epochs, batch 512, accumulation 1, context 64; later resume warning explicitly identifies saved world size 4 | 103,152 × 512 × 4 × 64 = **13,520,338,944** raw positions for this checkpoint trajectory |
|
| 43 |
+
| `ncp-owt-gpt2/checkpoint-iter-80000.pt` | Batch 96, accumulation 1, context 128, text prefix 64; log contains both 8- and 4-GPU resumes | Step 80k is not 50 completed epochs. At fixed 4/8 GPUs, 80k updates correspond to **3.932/7.864B** raw positions, half of which are target positions. Exact cumulative accounting requires resolving the resume timeline |
|
| 44 |
+
|
| 45 |
+
The generator's resolved planned totals include 1,375,350 and 2,750,699 updates
|
| 46 |
+
with `n_epochs=50`. The released artifact is only step 80,000. The logs include
|
| 47 |
+
restarts/resumes; discarded computation is separate from the retained training
|
| 48 |
+
trajectory. Do not report the configured endpoint as completed training.
|
| 49 |
+
|
| 50 |
+
The released tokenizer count implies about 2.704B corpus tokens, consistent
|
| 51 |
+
with the current local uint16 binary (2,704,046,552 tokens; rounding to whole
|
| 52 |
+
updates explains the small difference). This is a size consistency check,
|
| 53 |
+
**not a dataset-content/hash identity claim**.
|
| 54 |
+
|
| 55 |
+
Repository evaluation records additionally reference checkpoints up to 140k
|
| 56 |
+
(`results/gen_ppl_ncp_owt_gpt2_full.json`) and a run named `300ep` through 105k
|
| 57 |
+
(`results/gen_ppl_ncp_rope_qknorm_corrupt_300ep.json`). Names/configured epochs
|
| 58 |
+
are not proof of 300 completed epochs or of their exact token budgets.
|
| 59 |
+
|
| 60 |
+
## Current MsTok accounting
|
| 61 |
+
|
| 62 |
+
Batch 32/GPU × 12 accumulation × 8 GPUs × 256 = **786,432 positions/update**.
|
| 63 |
+
The current binary is 5,408,093,104 bytes of uint16 = 2,704,046,552 tokens.
|
| 64 |
+
|
| 65 |
+
| Updates per component | Repeated raw token positions | Nominal corpus passes |
|
| 66 |
+
|---|---:|---:|
|
| 67 |
+
| 10,000 | 7.864B | 2.91 |
|
| 68 |
+
| 17,192 | 13.520B | 5.00 |
|
| 69 |
+
| 25,000 | 19.661B | 7.27 |
|
| 70 |
+
| 34,384 | 27.041B | 10.00 |
|
| 71 |
+
|
| 72 |
+
The legacy 103,152-update tokenizer and our 17,192-update component therefore
|
| 73 |
+
have the **same raw-position budget**, despite a sixfold step-count difference.
|
| 74 |
+
At current batch/context, 131B positions require **166,576 updates**. At the
|
| 75 |
+
observed ~0.258 joint updates/second, that is roughly 7.5 days of training alone
|
| 76 |
+
on one eight-H100 node, excluding evaluations and overhead. More data is not
|
| 77 |
+
free and a multi-candidate sweep at that budget will not fit a ten-day deadline.
|
| 78 |
+
|
| 79 |
+
Before revising final-run lengths, identify which NCM checkpoint/result is the
|
| 80 |
+
comparison target. A short-budget tuning experiment remains useful, but is not
|
| 81 |
+
the same claim as matching a long-budget baseline.
|
README.md
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
tags:
|
| 5 |
+
- text-generation
|
| 6 |
+
- research
|
| 7 |
+
---
|
| 8 |
+
# ICLR debug experiments
|
| 9 |
+
|
| 10 |
+
Public artifacts for the `iclr-debug` branch of HazyResearch/MsTok.
|
| 11 |
+
W&B project: `mstok/iclr-debug`.
|
| 12 |
+
|
| 13 |
+
This repository currently establishes the experiment namespace and contains a
|
| 14 |
+
training-budget audit. No new model weights or final benchmark claims are
|
| 15 |
+
implied. Future selected checkpoints and evaluations will be organized by
|
| 16 |
+
unique run ID and optimizer step.
|
| 17 |
+
|
| 18 |
+
Training budgets must report raw token positions per component, dataset size,
|
| 19 |
+
global batch, optimizer updates, and LR horizons. Generation reports must label
|
| 20 |
+
supplied-level-zero versus unconditional sampling and record the reference
|
| 21 |
+
model precision, entropy, scored-token counts, and per-seed results.
|
| 22 |
+
|
| 23 |
+
See `NCM-BUDGET-AUDIT.md` for historical budget provenance and limitations.
|