ounce100m-code / docs /02-mix-plan.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
36 kB
# 02 β€” Mix plan: sources, counts, rationale, and the build
Phase 1 deliverable. Target: **1.0 B trained tokens** (Β§2 band 900M–1.1B; frozen number confirmed at
Gate 3). Every repo id, size, date and licence below was read live on 2026-09-19 from the Hub card, the
`/api/datasets/<id>` metadata endpoint, or parquet footers over HTTP β€” not from memory. Where a card and
the stored data disagree, that is recorded rather than smoothed over.
## 1. Selection rules applied, operationally
- **Β§3.4 English only.** Every source is either English-config-scoped (`20250702.en`, `eng_Latn`,
`sample/10BT`) or carries a per-row language field we filter on. Code and math are allowed and are
deliberately included.
- **Β§3.5 No raw crawls.** Excluded as a class: `allenai/c4`, `HuggingFaceFW/fineweb`,
`Skylion007/openwebtext`, `tiiuae/falcon-refinedweb`. What is admitted instead is
**quality-scored** (`fineweb-edu` carries a per-row classifier `score`/`int_score`),
**curated/rewritten** (`finewiki`, `cosmopedia`), **synthetic-textbook** (`finephrase`), or
**openly-licensed domain corpora** (`common-pile/*`, `stack-v3-train`). Common-Crawl appears only as
the *provenance* of a classifier-filtered subset, never as the selection criterion.
- **Β§3.3 No contamination.** Two mechanisms, not one assertion:
1. **Prefer sources with published decontamination.** `HuggingFaceTB/finemath` documents 13-gram
removal against GSM8K, MATH, MMLU and ARC test sets, with a public audit log
(`HuggingFaceTB/finemath_contamination_report`).
2. **Reject sources whose own cards validate on the eval suite.** That is why `openbmb/UltraData-Math`
is **dropped entirely** despite being attractive on paper (apache-2.0, 170B+ tokens): its card
lists MMLU/ARC-E/ARC-C/HellaSwag/PIQA/Winogrande/GSM8K/MATH500 as validation targets and its L3
tiers are synthetic *exam-shaped* text. Its en+zh mixing would also need a language filter.
Similarly excluded: `nvidia/Nemotron-Math-v2` (AoPS/SE solution traces, i.e. adjacent to GSM8K's
own source pool, and solution-formatted in a way that quietly steers toward instruction data).
3. The mechanical overlap audit of Β§5 runs against **validation/dev material only**; test splits stay
untouched until Phase 6.
## 2. The mix: 1.31 B raw β†’ ~1.05 B after dedup β†’ ~1.0 B trained
| # | Source (config) | raw tokens | GB to pull | licence | rationale, from corpus properties |
|---|---|---|---|---|---|
| 1 | `HuggingFaceFW/fineweb-edu` `sample/10BT`, keep `int_scoreβ‰₯4` | 300M | 1.1 | ODC-BY | Classifier-scored general English; the only cheap source of broad register coverage. Score-gating is a *quality* filter, which Β§3.5 permits, and it is the source's own documented axis. |
| 2 | `HuggingFaceFW/finepdfs-edu` `data/eng_Latn` | 150M | 0.63 | ODC-BY + CC ToU | Globally-deduplicated long-form educational PDFs. Long document structure teaches coherence over >1k tokens, which chunked web prose does not. `eng_Latn` only (69 scripts exist). |
| 3 | `HuggingFaceTB/cosmopedia` stanford+openstax+khanacademy+auto_math_text | 130M | 0.33 | **Apache-2.0** | Synthetic textbooks: causal and explanatory chains ("because", "therefore", worked steps). This is the register that builds science-QA and multi-step reasoning ability. Per-config GB measured, not estimated. |
| 4 | `HuggingFaceTB/cosmopedia` `wikihow` | 30M | 0.05 | Apache-2.0 | Procedural how-to steps β†’ ordered physical-world inference. Small slice, specific job. |
| 5 | `HuggingFaceFW/finewiki` `data/enwiki` | 80M | 0.43 | **CC-BY-SA-4.0** | Link-resolved, rewritten encyclopedic prose; definitional and comparative sentence forms that factual/recall-style prompts reward. Heaviest GB/token in the mix, hence capped. |
| 6 | `omarkamali/wikipedia-monthly` `20250702.en` | 60M | 0.03 | CC-BY-SA-4.0 | Cheapest encyclopedic tokens available (0.5 GB/Btok) and *current*; drop the `raw_mediawiki` column or it dominates the byte budget. |
| 7 | `HuggingFaceTB/finemath` `finemath-4plus` | 130M | 0.25 | ODC-BY | Best math density per GB (1.9) **and** the only candidate with a published eval-set decontamination report. |
| 8 | `open-web-math/open-web-math` | 70M | 0.15 | ODC-BY (card body) | Forum/webbook math in natural prose, incl. LaTeX. Complements #7's textbook skew with multi-turn human reasoning text. Oldest source here (2023-10) β€” cited as such. |
| 9 | `HuggingFaceFW/finephrase` `tutorial` + `faq` | 90M | 0.31 | ODC-BY | Rewritten into stepwise-tutorial and question/answer registers β€” forms largely absent from organic prose. |
| 10 | `HuggingFaceFW/finephrase` `table` | 30M | 0.10 | ODC-BY | Tabular→narrative realisation: reading quantities, units and row/column structure. |
| 11 | `HuggingFaceCode/stack-v3-train`, `license_type=="permissive"`, per-language cap | 130M | 0.52 | ODC-BY + per-file SPDX | Code for compositional syntax and deterministic multi-step logic. **Must** filter to permissive licences (constraint 4) and cap XML/HTML/JSON/JS, which the card notes are byte-dominant. |
| 12 | `SimpleStories/SimpleStories` | 40M | 0.07 | MIT | Syntactically simple long-range narrative. Disproportionately useful below ~200M params, where the model otherwise never sees a consistent referent across a full context window. |
| 13 | `common-pile/libretexts` + `common-pile/arxiv_abstracts` | 40M | 0.08 | CC0 / per-doc PD/CC-BY | Openly-licensed OER textbooks and scientific abstracts; cleanest licence hygiene in the pool (arXiv metadata is CC0). |
| | **total raw** | **1,310M** | **β‰ˆ4.35 GB** | | |
**Headroom, and why it is sized like this.** Β§Phase 2 requires slack above the target for dedup loss
and a held-out split. Cross-source near-dedup is expected to cost **12–18 %** (FinePhrase's four configs
are rewrites of the *same* ~338.7M source documents, so they must be deduplicated against the shared
`id`/`url` or they buy topic echo rather than knowledge; #2/#5/#6 overlap on encyclopedic topics;
#1/#2 overlap on educational web content). 1,310M Γ— 0.85 β‰ˆ 1,114M, minus a 2 % validation holdout
(β‰ˆ22M) β†’ **~1.09 B available**, which brackets the 1.0 B target and leaves room to land at 0.9 B
without re-sourcing if Gate 3's throughput measurement forces the pre-registered fallback.
**Sits comfortably inside the instance.** β‰ˆ4.35 GB of shards + ~2 GB of Arrow/tokeniser working
overhead + 2 GB of final token store β‰ͺ the measured **19.5 GB** `/kaggle/working`. Every source here is
pullable as individual 270 MB–3 GB shards, so the working directory is never the binding constraint β€”
provided the build pulls shard-by-shard and deletes after tokenising, which Β§4 requires it to do.
## 3. Excluded, with reasons (so nobody re-litigates from scratch)
| excluded | why |
|---|---|
| `allenai/c4`, `HuggingFaceFW/fineweb`, `Skylion007/openwebtext`, `tiiuae/falcon-refinedweb` | raw/unfiltered crawls β†’ Β§3.5 |
| `openai/gsm8k`, `cais/mmlu`, `allenai/ai2_arc`, `hellaswag`, `piqa`, `winogrande`, `truthfulqa` | the eval tasks themselves β†’ Β§3.3. Also: never glob a directory that could catch these. |
| `nvidia/Nemotron-CC*`, `-Math-v1`, `-Code-v1` | `gated: manual`; anonymous card and file reads return *"Access to dataset … is restricted"* β†’ token counts unverifiable and republication status unclear |
| `openbmb/UltraData-Math` | card validates on the eval suite; exam-shaped L3 tiers; en+zh mixed |
| `nvidia/Nemotron-Math-v2` | solution-formatted AoPS/SE traces, GSM8K-adjacent; reads as instruction data |
| `allenai/peS2o` | ~11 GB per Btok by measurement β€” worst ratio in the pool for 42 B tokens we do not need |
| `HuggingFaceTB/smollm-corpus` `cosmopedia-v2` | duplicate role with #3; also a documented card-vs-data mismatch (card 17.8 B vs ~31.7 B implied by stored `token_length`) |
| `HuggingFaceFW/clean-wikipedia` | exists but README body is literally "Please see FineWiki instead" β†’ deprecated |
| `HuggingFaceTB/smollm-corpus` `python-edu` | **not a corpus.** Verified schema is `blob_id, repo_name, path, length_bytes, score, int_score` β€” a filter list pointing at gated `bigcode/the-stack-v2`, with no `text` column. Do not budget code tokens from it. |
| `bigcode/the-stack-v2-dedup`, `codeparrot/github-code` | not verified this session; `the-stack-smol` is `gated: auto` |
| `TinyLlama` 32k tokenizer | Llama-2-derived; redistribution licence status unverified β†’ excluded from both tokenizer and any data |
| `Cion-lab/Pretraining-10B-v1` (pre-existing in the account) | composition/mix unknown, token counts estimated with the **Qwen2.5** tokenizer, provenance not inspectable β†’ cannot be cleared under Β§3.5 or Β§3.3 |
## 4. Card-vs-data warning
Two sources disagree with themselves: `cosmopedia-v2` (17.8 B card vs ~31.7 B stored) and `finephrase`
(486 B "completion tokens" card vs ~1.4 B implied by stored rows Γ— inherited `token_count`). FinePhrase's
`token_count`/`score` columns describe the **source document, not the rewrite**. **Therefore: no token
number in this project is ever taken from a card.** Counts come from running our own tokenizer over the
bytes, which is also what the Gate 2 manifest will record.
## 5. Build pipeline (Phase 2 design, on Kaggle CPU)
1. **Pull shard-by-shard, delete after consuming.** 73–89 MB/s measured (35–89 across runs), so 4.35 GB
arrives in minutes; the constraint is *peak residency*, not bandwidth. Never materialise the mix.
2. **Normalise + filter.** Language filter on the row-level field where present; `int_scoreβ‰₯4` for
fineweb-edu; `license_type=="permissive"` + per-language cap for stack-v3; drop `raw_mediawiki`.
3. **Dedup, memory-bounded.** Exact: 64-bit hash of the normalised document. Near: MinHash + LSH over
13-gram shingles (matching finemath's documented practice), **processed in bands against a
disk-backed index**, because the box is a 30 GiB cgroup with **no swap** β€” a billion sketches in one
dict is how a free-CPU job dies at hour six.
4. **Contamination audit (Β§3.3).** Mechanical 13-gram overlap of the surviving mix against benchmark
**validation/dev** material only β€” including **MMLU's `dev` split specifically**, since those are the
few-shot exemplars the harness will use, so overlap there is contamination by construction. A script
computes and writes counts per source; **no item text is read, displayed or summarised by the agent.**
Any source failing a stated threshold is dropped or cleaned; result written up in `docs/`. **Test
splits are not opened until Phase 6.**
5. **Tokenise** with the frozen SmolLM2 BPE (49,152 β€” Β§3 of `01-plan`). Measured cost on this shape:
660,943 tok/s across 4 cores β†’ **1.31 B tokens in ~0.55 h**. Not a bottleneck; process-shard rather
than thread-shard, since rayon scaling saturated at 2.7Γ— on 4 cores.
6. **Shard for exact positional resumption (Β§Phase 2 requirement).** Concatenate token ids into fixed-size
**`uint16`** shards (`vocab 49,152 < 65,536`, so uint16 is exact and halves the bytes): **40 M tokens =
80 MB per shard**, ~28 shards, ~2 GB total. Then a resume position is literally
`(shard_index, token_offset)` β€” one integer pair, memmap'd, so `skip_first_batches` on a replaying
`Trainer` becomes an index advance rather than a decode. This is the property that makes Β§3.1's
"same data, same position" checkable instead of hoped for.
7. **Hold out validation:** ~22 M tokens sampled **proportionally across sources**, never whole documents
(whole-document holdout leaks topic statistics and makes PPL optimistic), written as its own shard so
validation PPL is bit-reproducible at Gate 4.
8. **Publish** as public `ounce100m-mix-v1` with a manifest recording: per-source and total token counts,
shard count and size, tokenizer identity (repo + revision + sha256 of `tokenizer.json`), per-shard
sha256, the dedup parameters, the audit counts, and the cursor semantics of step 6.
## 6. Licensing the derived mix
#5 and #6 are **CC-BY-SA-4.0** (share-alike), and `wikimedia` additionally carries GFDL. A published
derivative containing them must be offered under SA with attribution and a change log. That is compatible
with Β§2's "repos are public", so the plan is: **license the mix CC-BY-SA-4.0** and ship an
`ATTRIBUTION.md` listing every source, its licence, and the transformations applied. ODC-BY sources
(#1, #2, #7–#10) require notices kept and changes documented; `stack-v3-train` requires the per-file
SPDX filter of Β§2 step 11 so we never republish a non-permissive file.
**If a fully permissive mix is later required**, drop #5 and #6 (140 M raw) and backfill from #1/#3. The
trade is a small loss of encyclopedic density for a licence with no reciprocal obligation β€” recorded
here so the choice is available and deliberate rather than discovered under time pressure.
## 7. Gate 2 checklist, restated as things that must be true
- [ ] dataset live on the Hub as a public `ounce100m-*` repo
- [ ] token count **verified against the manifest** by an independent recount, not by the builder's own number
- [ ] contamination audit passed **and written up** (per-source overlap counts, validation/dev only)
- [ ] validation split reserved, its token count recorded
- [ ] shard layout proven small enough that a training job reads a few shards at a time and never
materialises the mix β€” demonstrated by a job that reports peak `/kaggle/working` usage while
consuming shards, not asserted from the arithmetic above
- [ ] exact-position cursor `(shard_index, token_offset)` demonstrated to be reproducible
## 8. Risks
1. **Dedup is the only stage with real wall-clock cost** and the only one that can exhaust RAM. Design
for bands + disk; measure on one source before all thirteen.
2. **Unauthenticated Hub rate limits.** Jobs currently read anonymously and the API warns about it; at
~50 shard downloads that is probably fine, but 429s mid-build are plausible. Mitigation is the
mounted-credential path (Β§8 of `01-plan`), which is being verified now.
3. **Card token counts are unreliable** (Β§4) β€” every number in this file is either measured over HTTP or
will be replaced by our own tokeniser's count at Gate 2.
4. **`finepdfs-edu` adds Common Crawl's Terms of Use** on top of ODC-BY. Included because the subset is
classifier-selected and globally deduplicated rather than a dump, but it is the least clean licence
admission in the mix and the easiest to cut if the attribution review gets nervous.
5. **Two sources are rewrites/synthetic** (#3, #4, #9, #10). That is permitted explicitly by Β§3.5, but a
34 % synthetic share can produce fluent-but-homogeneous prose; the proportions above keep organic
English (sources 1, 2, 5, 6, 8, 12, 13) in the majority, and that balance is a deliberate choice
worth revisiting only with evidence.
---
## 9. Phase 2 step 1 ran β€” Β§2 corrected by measurement
`dodosoomro/ounce100m-p2-source-inventory` (CPU, 116 s, 1,500 rows per source, tokenised with the
**frozen SmolLM2 tokenizer**: vocab 49,152 confirmed, decode round-trips, 5.88 chars/token on a formal
probe sentence). **13 of 18 entries resolved; 5 failed on config names, and two sources turned out to be
unusable as described.** Β§2's estimates were right in aggregate (predicted β‰ˆ4.35 GB, measured 4.46 GB for
the 570 M tokens that resolved) but wrong in specifics, so the specifics are replaced rather than
averaged.
### 9.1 Config-name corrections (the datasets-server truth, from the error messages themselves)
| Β§2 said | correct value | note |
|---|---|---|
| `fineweb-edu` `sample/10BT` | **`sample-10BT`** | hyphen, not a path. Full set: `default`, `sample-10BT`, `sample-100BT`, `sample-350BT`, plus per-dump `CC-MAIN-2025-05/08/13/18/21/26…` |
| `finemath` `finemath4plus` | **`finemath-4plus`** | also `finemath-3plus`, `infiwebmath-3plus`, `infiwebmath-4plus` |
| `finewiki` `data/enwiki` | config **`en`** | configs are plain language codes (`ab`, `ace`, `af`, …), not directory paths |
| `wikipedia-monthly` `20250702.en` | **`20260101.en`** | the 2025-07 config is gone; **newer** dumps now exist (`20260101.*`), which is better for currency anyway |
| `stack-v3-train` `python` | **`default`** only | **there are no per-language configs.** Language is a row field, and rows are whole repositories with a nested `files` struct |
Every one of these would have been a mid-build failure. Recorded in `docs/00-platform-notes.md` style:
the card's own naming is not an API β€” `load_dataset`'s error is.
### 9.2 Measured yield per document β€” this is what the token arithmetic actually runs on
| source | tokens/doc | chars/doc | chars/token | non-English-flagged rows | verdict |
|---|---|---|---|---|---|
| `finepdfs-edu` `eng_Latn` | **1,321** | 14,352 | 3.98 | **128/1500 = 8.5 %** | keep, **must add our own language filter** |
| `cosmopedia` `stanford` | 932 | 4,917 | 5.30 | 15/1500 = 1.0 % | keep |
| `cosmopedia` `wikihow` | 907 | 4,541 | 4.97 | 0.7 % | keep |
| `cosmopedia` `openstax` | 765 | 3,820 | 4.93 | 0.8 % | keep |
| `cosmopedia` `auto_math_text` | 689 | 2,811 | 4.14 | 0.8 % | keep |
| `cosmopedia` `khanacademy` | 677 | 3,047 | 4.26 | 1.1 % | keep |
| `open-web-math` `default` | **1,182** | 6,795 | **3.31** | **112/1500 = 7.5 %** | keep, needs en filter; low chars/token = LaTeX density, which is the point |
| `finephrase` `tutorial` | 779 | 4,903 | 4.58 | 1.1 % | keep. Column is **`text`**, not `completion` |
| `finephrase` `faq` | 744 | 4,965 | 4.60 | 1.1 % | keep |
| `finephrase` `table` | 715 | 4,772 | 4.59 | 1.1 % | keep |
| `SimpleStories` | **285** | 1,260 | 4.33 | 0.5 % | keep, but it is 4Γ— the *row* count for the same tokens |
| `common-pile/arxiv_abstracts` | 185 | 791 | 4.31 | 2.2 % | keep. Column is **`text`**, not `raw_content` |
| `common-pile/libretexts` | **75** | 964 | **2.42** | **1396/1500 = 93 %** | **DROP** |
**`libretexts` is disqualified, and the reason matters.** Its `text` field is not prose: 93 % of sampled
rows failed a naive English test, chars/token collapsed to 2.42, and the first example is a bare
`https://math.libretexts.org/Bookshelves/...` URL. The Common-Pile schema puts real content in
`metadata`/`raw_content` behind a per-source convention, so the column we assumed is an identifier field.
Rather than reverse-engineer four sub-corpora's schemas for 10–20M tokens of marginal value, it is cut.
`arxiv_abstracts` from the same publisher *did* work (its `text` is real abstract prose), so this is a
per-dataset fact and not a Common-Pile-wide one.
### 9.3 Two open problems this created, and how they are being handled
1. **Code no longer has a cheap path.** With no per-language config, getting 60–130M tokens of Python/JS
from `stack-v3-train` means streaming whole-repo rows, opening a nested `files` struct, and filtering
on a row field β€” for a source that concentrates in XML/HTML/JSON anyway. Code is not mandatory (Β§3.4
makes proportions our call), and at 100M params / 1B tokens the marginal value is genuinely unclear.
**Action: a dedicated mini-probe before committing** β€” measure rows-sampled-per-Python-repo and the
nested schema β€” and hold code at **≀60M** in the plan until that returns. If it is awkward, drop it;
that is a defensible Β§3.4 choice, not a failure.
2. **`eng_Latn` and `open-web-math` are not actually English-only** (8.5 % and 7.5 % of sampled rows fail
a crude English test; the finepdfs sample included Cyrillic inline in a single "English" document).
Β§3.4 is English-only, so **a language filter is now a required pipeline stage, not a courtesy** β€”
applied per row on our side, on top of whatever the upstream config claims. `finepdfs` carries
`full_doc_lid` / `per_page_languages` columns to do this cheaply; the general fallback is
`language_score` where present and a fastText-style gate where not.
### 9.4 Revised totals
Confirmed-and-usable now: 570M target tokens across 13 resolved configs, **β‰ˆ4.5 GB**, at measured yields
of roughly **909k documents** β€” plus `fineweb-edu`/`finewiki`/`wikipedia-monthly`/`finemath` recoverable
with the corrected config names (Β§9.1), which restores ~570M of the planned 1,310M. Net effect on the
plan: **the shape survives, the line items do not.** Β§2's per-source token targets stay as targets;
Β§2's config strings and two column names are superseded by this section.
**Directly measured consequence for the build:** 909k documents for 570M tokens β‰ˆ 627 tokens/document
average, so the mix is ~1.6–2.1M documents for 1.0–1.3B tokens. That is the number the dedup design has
to be sized against β€” under 2.1M documents, a MinHash signature table comfortably fits the 30 GiB RAM,
which materially de-risks the one stage Β§8 called the real cost.
## 10. Builder implemented, and one blocking flaw found and fixed
`code/build/build_mix.py` (mirrored in `Cion-lab/ounce100m-code`) is the Phase 2 builder. It stages each
source separately (stream β†’ normalise β†’ gates β†’ dedup β†’ tokenise) and then **merges** the staged shards
into the final mixed shards by largest-deficit round-robin over documents.
**E-016, and why the smoke test was mandatory.** The first version consumed one source to its target and
only flushed a shard when the buffer filled. Since fineweb-edu's target alone is ~37 shards of 8M tokens,
the first third of the mix would have been *nothing but web text*, then textbooks, then math β€” a curriculum
drift across the run that no LR schedule can repair, and one the builder itself reported as success
(checksums, headers, token counts and cursor probes were all clean). The verifier caught it as
`distinct_sources_per_shard_min: 1`. The shuffle that was supposed to fix ordering only permuted documents
*inside* a single-source buffer. After the two-stage rewrite, the same smoke run reports
**`min = max = median = 4`** sources in every shard, 12,001,386 tokens recounted from the bytes,
0 out-of-range ids, 40/40 cursor probes and 40/40 document-boundary reconstructions correct.
**Shard format** (shared by staged and final, so document boundaries survive to the trainer):
```
header : <IIII = vocab, n_docs, n_tokens, flags 16 B
offsets : uint32 Γ— (n_docs + 1), cumulative, [0] = 0
ids : uint16 Γ— n_tokens (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes)
doc i = ids[offsets[i] : offsets[i+1]]
```
Target 8M tokens/shard (16 MB, ~150 shards, ~2.5 GB for the mix). The resume cursor is
`(shard_index, token_offset)` β€” one integer pair, which is what makes Β§3.1's "same data, same position"
checkable and what keeps `Trainer`'s replaying `skip_first_batches` cheap on a map-style read.
**Interruption safety, which was not optional.** `/kaggle/working` is destroyed when a session ends, and a
full staging pass is hours of streaming. So each source is published to `Cion-lab/ounce100m-mix-stage`
(the *stage* repo, distinct from the final `ounce100m-mix-v1`) the moment it completes, with a
`record.json` holding its shard list, token count and drop statistics. A new session rebuilds its stage
bookkeeping from the Hub and skips every source already published; only the source in flight at the moment
of the kill is re-staged. `code/build/smoke_hub_resume.py` rehearses this by wiping the local tree between
stage and merge, because "merge-only with an empty state and no restore path" is a failure mode that exits
with code 0 while emitting an empty mix.
**Revised raw total: 1.23 B tokens = 23 % headroom** over the 1.0 B trained target (raised fineweb-edu to
340M, finepdfs-edu to 180M, finemath to 160M, because Β§9's target edits had quietly pulled the plan from
1.31 B down to ~1.13 B β€” about 11 % headroom, which is not headroom once a 12–18 % dedup loss is applied).
16 sources, no code source yet (Β§9.3).
**Gates the build must satisfy before Gate 2 is claimed**, as assertions in `verify_mix.py` rather than
prose: every shard sha256 matches the manifest; no trailing bytes; offsets well-formed; zero out-of-range
ids; total tokens recounted from bytes equals the manifest; and **β‰₯ half the distinct sources present in
every shard**. The last one exists because the first build passed everything else.
## 11. Interruption safety proven, and the build that this authorises
A CPU session's `/kaggle/working` is destroyed when it ends. Everything below was measured, not assumed
(`dodosoomro/ounce100m-p2-cold-resume-rehearsal` **v6**, 17:35Z, 55 s wall-clock, pinned REV `428c5234`).
| Step | Observed |
|---|---|
| 1. stage two sources, publish each | 2,400,867 tokens staged; `record.json` + `shard-0000.bin` per source on the Hub |
| 2. **wipe the local tree** (`root_exists: false`, `stage_dir_exists: false`) | faithful simulation of a session kill, not a soft restart |
| 3. `--merge-only` with the Hub as the only state | `hub: restored 2 staged source(s) (2,400,867 tokens)` β†’ merged shard 2,352,747 tok / 5,815 docs, **2/2 sources in it**, val shard 48,120 tok |
| 4. `verify_mix.py` over the restored mix | `PASS: true`, `sha_ok 1/1`, `out_of_range_ids: 0`, `trailing_bytes: 0`, cursor probes **40/40 boundaries reconstructed, 0 mismatches**, `recount_matches_manifest: true` |
| 5. verdict + self-cleanup | `COLD_RESUME_PASSED=True`, `cleanup=deleted`, throwaway repo confirmed gone from the namespace |
v4 of the same rehearsal had reported `false` over an identically-working pipeline: a stray
`print("shards=", …)` inside the `BUILD_JSON` block made the parent's `json.loads` fail and a helper fold
that into `None` (memory/ERRORS.md **E-019**, and the report-shape recurrence **E-023**). The fix was to
the harness, plus a `(obj, why_null)` return and a `manifest.json` fallback so a verdict can never again
be silently null.
**The build now running** (`dodosoomro/ounce100m-p2-full-mix-build`, 17:38Z, CPU, `sessionTimeoutSeconds
43200`, pinned REV `98e49560`): `build_mix.py --root /kaggle/working/mixroot --hub-repo
Cion-lab/ounce100m-mix-stage` with defaults β€” stage cap 1.31 B tokens, per-source targets summing to
1.23 B, 8 M tokens/mixed shard, `VAL_STRIDE 50` holdout, shuffle seed 20260919. Each source publishes as
it completes, so an interruption costs at most one source's streaming time, and a re-launch restores the
finished ones from the Hub (step 3 above is that exact path).
**Audit semantics, hardened the same way.** `audit_contamination.py` compares 13-token windows against
`{train, validation, dev}` of the eight tasks β€” never `test` β€” and emits counts only. It now also reports
`shards_missing`, and its verdict is composite: `AUDIT_PASSED` requires every manifest shard present and
read, > 0 documents and grams measured, β‰₯ 6 readable reference sets, zero reference errors, and zero
overlap. Exit 6 means *overlap found* (a real finding); exit 7 means *the audit did not measure* (a bug).
A zero-overlap report over shards that were not on disk is no longer expressible.
**Publishing ships the tokenizer.** `publish_mix.py` copies `tokenizer.json` into the dataset root,
re-hashes it and aborts with exit 8 unless it matches the sha256 recorded in the manifest β€” so the
published ids and the published tokenizer are provably the pair that produced them (Β§3.12).
## 12. Gate 2 β€” the record, including where it failed and what that failure was worth
**Merge and verify: passed, twice, reproducibly.** From the 15 staged sources
(`Cion-lab/ounce100m-mix-stage`, 1,150,015,841 tokens staged):
| field | value |
|---|---|
| merged shards | **141** (each drawn from all 15 sources) |
| train tokens | **1,127,163,952** |
| held-out validation | **22,851,889** in 12 shards (every 50th merged document) |
| documents | 1,078,807 |
| per-source share | 75,144,263 tokens each, equal by construction |
| merge time | 16.4 s (restore from the Hub included: 123.8 s) |
| `verify_mix` | **`PASS: true`** β€” sha256 141/141, headers 141/141, 0 bad offsets, 0 out-of-range ids, 0 trailing bytes, recount matches manifest, 40/40 cursor probes, `distinct_sources_per_shard_min = 15` (floor was 7) |
| `finewiki__en` | **never staged.** Gate 2's own diagnostic showed the columns present and `first_row_ok: true` with a truncated `ImportError` in the loading script. 15 sources at 1.127 B is above the 1.0 B target with 12.7 % headroom, so this is a composition note, not a blocker |
**Then the audit refused to publish, and it was right to.** `audit_contamination` exited 7:
| task | 13-gram hits | reference size (verified against the dataset server) |
|---|---|---|
| mmlu | 904,288 | dev 285 + validation 1,531 items β‰ˆ 149k grams |
| gsm8k | 185,600 | train 7,473 items = 360,656 grams |
| hellaswag | 1,852 | train 39,905 + validation 10,042 |
| arc_easy | 358 | train 2,251 + validation 570 |
| arc_challenge | 214 | train 1,119 + validation 299 |
| piqa | **not measured** | the loader yielded 0 of 16,113 train rows without raising |
| winogrande | (not in the visible excerpt) | train 40,398 + validation 1,267 |
| truthfulqa | (repo name was wrong) | validation 817, `multiple_choice` |
`expected_false_positives` for the whole scan was **β‰ˆ0.0001**, so these are matches, not collisions. The
first reading I wrote in `STATE.md` β€” "GSM8K's whole train split is in the mix" β€” was an over-claim: hits
are counted per *mix* gram, so a small number of frequently repeated reference grams can produce a large
count, and 185,600 hits against 360,656 distinct GSM8K grams means "up to half of its gram vocabulary
occurs somewhere in 1.1 B tokens". The columns that settle severity are `mix_documents_majority_hit`
(documents whose grams are more than half matched β†’ verbatim inclusion) and `source_hit_document_fraction`
(which source, and how much of it) β€” and both are only computable on the fixed audit, which is what Gate
2c/2d exist to produce.
**Why upstream did not catch it, verified rather than assumed:** FineMath's card states it removed 13-gram
overlap against the **test** sets of GSM8k, MATH, MMLU and ARC; Cosmopedia's describes a 10-gram +
`difflib` pass against **test benchmarks**; OpenWebMath documents only SimHash self-deduplication with no
benchmark claim. `validation`/`dev`/`train` overlap was therefore never in scope for any of them β€” and
`validation` is the split lm-eval scores HellaSwag, PIQA, Winogrande and TruthfulQA on
(`docs/05-eval-plan.md` Β§7).
**Gate 2 remained unsigned at the end of this stage.** `Cion-lab/ounce100m-mix-v1` did not exist yet. The
path to it: attribution β†’
drop the offending documents by mask (or a source if attribution says it is broadly soaked) β†’ re-merge
with `--exclude-dir` β†’ re-audit until `AUDIT_PASSED` with all 8 tasks covered β†’ `publish_mix`, which now
refuses to ship unless the audit passed, covered all 8, scanned the val shards, and its `mix_bytes_sha256`
matches the exact manifest being published. `E-030` explains why the first filtered re-merge could not
work and what was changed.
### 12.4 Resolution β€” the filtered mix, measured (Gates 2b–2f, 2026-09-19)
**Attribution first, always.** Gate 2b/2c scanned the *staged* shards per source, which is the only way to
say where contamination lives rather than how much the merged mix contains. Of 1,955,832 staged documents,
**10,425 (0.533 %)** carried at least one matched 13-gram and **0 were majority-matched** β€” shared spans,
not wholesale inclusion. The count concentrated in five sources: finemath-4plus 2,688 documents, fineweb-edu
1,043, cosmopedia `auto_math_text` 568, finepdfs-edu 420, open-web-math 339, with the remaining ~5,367
spread over ten more. Per-source *hit counts* summed to ~50,700 while the merged scan reported 2,096,540 for
essentially the same gram volume; the discrepancy is unreconciled and recorded as **E-032c** rather than
smoothed over. It changes no decision: the exclusion is per-document and the post-filter overlap is zero
either way.
**The exclusion, and what it cost.** `audit_contamination --write-filter` emitted one `.u8` mask per source
(1 byte per staged document, ~700 KB total) marking documents that contain a matched window. Re-merging with
`--exclude-dir` dropped **5,878 documents = 17,366,967 tokens** (1.5 % of the mix) and the re-audit of the
result gave **`overlap_total = 0`** with 8/8 benchmark tasks covered, the held-out validation shards
scanned clean, and `MEASURED: true`. Final mix: **1,109,714,831 training tokens** in 141 shards plus
22,851,889 held-out validation tokens β€” still **11.0 %** above the 1.0 B target, so the pre-registered
0.9 B fallback (D-012) did not fire and the contamination fix cost nothing that the token budget had to
pay for.
**Reference coverage is now a precondition, not an assumption.** Gate 2b found that `piqa`'s train and
validation splits returned **0 rows** from the `datasets` library inside the Kaggle image while the server
reports 16,113 + 1,838 β€” the first audit had scored that as "clean". The audit now fetches each split's
server-side row count (`/size`), refuses any split it read to under 90 % of it, retries through the HTTP
`/rows` endpoint, and marks the whole report `MEASURED: false` if a loader still falls short. TruthfulQA's
answer text also sits under keys (`mc1_targets`, `mc2_targets`, `abstractions`) the first version never
read. All reference splits used are train/validation/dev; **no test split was opened** (Β§3.3).
**Why publication took three attempts (E-033).** Gates 2e and 2f both built and audited the mix correctly
and then died in the uploader: `publish_mix` sent one commit per file, and the Hub began rejecting every
request β€” including 30-byte files β€” after ~139 commits. That is a commit-rate ceiling, not a network fault,
and the "resumable" re-send was itself broken because `RepoFile.lfs` is a dict and `getattr(ent.lfs,
"sha256", None)` is therefore `None` for every shard, i.e. *nothing* looked published. The uploader now
reads `lfs["sha256"] or lfs["oid"]`, sends the remaining files as **one `upload_folder` commit**, and after
the commit re-lists the repo and re-hashes, exiting 11 if a path is absent and 12 if a remote digest
disagrees with the local bytes. `Cion-lab/ounce100m-mix-v1` held 142 of 173 files when Gate 2g re-ran the
same pipeline with it.
**Gate 2 is signed** (`dodosoomro/ounce100m-gate-2g-publish-mix-v1`, complete 2026-09-20T00:31Z, CPU, zero
GPU quota). The rebuilt mix reproduced Gate 2e's token total exactly β€” **1,109,714,831** β€” which is the
determinism claim in Β§11 tested a second time on a different session with the same inputs. `audit.json` on
the Hub reports `overlap_total = 0`, `mix_documents_with_any_hit = 0`, `majority_hit_docs = 0`,
`expected_false_positives = 0.0`, `tasks_covered` = all eight, `val_hits = 0` across 21,898 held-out
documents, and `AUDIT_PASSED`, `MEASURED`, `COVERED_ALL_SHARDS` all true, bound to these exact bytes by
`mix_bytes_sha256 = 732eec7f19e77f36…`. `verify_mix` exited 0. The batched uploader landed all 174 files
(139 train shards, 12 val shards, 16 `filter/` objects, `manifest.json`, `audit.json`, `tokenizer.json`,
`README.md`, `ATTRIBUTION.md`, `build_state.json`, `.gitattributes`), and four train plus two validation
shards were then read back from the anonymous `resolve` endpoint and matched the manifest's sha256 β€” so the
mix is public and serving, not merely listed (Β§3.13).
One editorial note on this section: Β§12.2's first draft treated `overlap_total` as a count of contaminated
documents. It is a count of matched *grams*, which is why a 41Γ— difference between the per-source and
merged hit totals could coexist with a 1.3 % agreement on document counts (E-032c). The action taken β€”
drop the documents, re-audit to zero β€” was unaffected, but the number in the record was read wrongly for
two gates, and the fix was to write the density columns (`mix_documents_with_any_hit`,
`mix_documents_majority_hit`) and quote *those* when judging severity.