|
Download docs/02-mix-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 36 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/02-mix-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/02-mix-plan.md
-
curl -L -o 02-mix-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/02-mix-plan.md
36 kB
| # 02 β Mix plan: sources, counts, rationale, and the build | |
| Phase 1 deliverable. Target: **1.0 B trained tokens** (Β§2 band 900Mβ1.1B; frozen number confirmed at | |
| Gate 3). Every repo id, size, date and licence below was read live on 2026-09-19 from the Hub card, the | |
| `/api/datasets/<id>` metadata endpoint, or parquet footers over HTTP β not from memory. Where a card and | |
| the stored data disagree, that is recorded rather than smoothed over. | |
| ## 1. Selection rules applied, operationally | |
| - **Β§3.4 English only.** Every source is either English-config-scoped (`20250702.en`, `eng_Latn`, | |
| `sample/10BT`) or carries a per-row language field we filter on. Code and math are allowed and are | |
| deliberately included. | |
| - **Β§3.5 No raw crawls.** Excluded as a class: `allenai/c4`, `HuggingFaceFW/fineweb`, | |
| `Skylion007/openwebtext`, `tiiuae/falcon-refinedweb`. What is admitted instead is | |
| **quality-scored** (`fineweb-edu` carries a per-row classifier `score`/`int_score`), | |
| **curated/rewritten** (`finewiki`, `cosmopedia`), **synthetic-textbook** (`finephrase`), or | |
| **openly-licensed domain corpora** (`common-pile/*`, `stack-v3-train`). Common-Crawl appears only as | |
| the *provenance* of a classifier-filtered subset, never as the selection criterion. | |
| - **Β§3.3 No contamination.** Two mechanisms, not one assertion: | |
| 1. **Prefer sources with published decontamination.** `HuggingFaceTB/finemath` documents 13-gram | |
| removal against GSM8K, MATH, MMLU and ARC test sets, with a public audit log | |
| (`HuggingFaceTB/finemath_contamination_report`). | |
| 2. **Reject sources whose own cards validate on the eval suite.** That is why `openbmb/UltraData-Math` | |
| is **dropped entirely** despite being attractive on paper (apache-2.0, 170B+ tokens): its card | |
| lists MMLU/ARC-E/ARC-C/HellaSwag/PIQA/Winogrande/GSM8K/MATH500 as validation targets and its L3 | |
| tiers are synthetic *exam-shaped* text. Its en+zh mixing would also need a language filter. | |
| Similarly excluded: `nvidia/Nemotron-Math-v2` (AoPS/SE solution traces, i.e. adjacent to GSM8K's | |
| own source pool, and solution-formatted in a way that quietly steers toward instruction data). | |
| 3. The mechanical overlap audit of Β§5 runs against **validation/dev material only**; test splits stay | |
| untouched until Phase 6. | |
| ## 2. The mix: 1.31 B raw β ~1.05 B after dedup β ~1.0 B trained | |
| | # | Source (config) | raw tokens | GB to pull | licence | rationale, from corpus properties | | |
| |---|---|---|---|---|---| | |
| | 1 | `HuggingFaceFW/fineweb-edu` `sample/10BT`, keep `int_scoreβ₯4` | 300M | 1.1 | ODC-BY | Classifier-scored general English; the only cheap source of broad register coverage. Score-gating is a *quality* filter, which Β§3.5 permits, and it is the source's own documented axis. | | |
| | 2 | `HuggingFaceFW/finepdfs-edu` `data/eng_Latn` | 150M | 0.63 | ODC-BY + CC ToU | Globally-deduplicated long-form educational PDFs. Long document structure teaches coherence over >1k tokens, which chunked web prose does not. `eng_Latn` only (69 scripts exist). | | |
| | 3 | `HuggingFaceTB/cosmopedia` stanford+openstax+khanacademy+auto_math_text | 130M | 0.33 | **Apache-2.0** | Synthetic textbooks: causal and explanatory chains ("because", "therefore", worked steps). This is the register that builds science-QA and multi-step reasoning ability. Per-config GB measured, not estimated. | | |
| | 4 | `HuggingFaceTB/cosmopedia` `wikihow` | 30M | 0.05 | Apache-2.0 | Procedural how-to steps β ordered physical-world inference. Small slice, specific job. | | |
| | 5 | `HuggingFaceFW/finewiki` `data/enwiki` | 80M | 0.43 | **CC-BY-SA-4.0** | Link-resolved, rewritten encyclopedic prose; definitional and comparative sentence forms that factual/recall-style prompts reward. Heaviest GB/token in the mix, hence capped. | | |
| | 6 | `omarkamali/wikipedia-monthly` `20250702.en` | 60M | 0.03 | CC-BY-SA-4.0 | Cheapest encyclopedic tokens available (0.5 GB/Btok) and *current*; drop the `raw_mediawiki` column or it dominates the byte budget. | | |
| | 7 | `HuggingFaceTB/finemath` `finemath-4plus` | 130M | 0.25 | ODC-BY | Best math density per GB (1.9) **and** the only candidate with a published eval-set decontamination report. | | |
| | 8 | `open-web-math/open-web-math` | 70M | 0.15 | ODC-BY (card body) | Forum/webbook math in natural prose, incl. LaTeX. Complements #7's textbook skew with multi-turn human reasoning text. Oldest source here (2023-10) β cited as such. | | |
| | 9 | `HuggingFaceFW/finephrase` `tutorial` + `faq` | 90M | 0.31 | ODC-BY | Rewritten into stepwise-tutorial and question/answer registers β forms largely absent from organic prose. | | |
| | 10 | `HuggingFaceFW/finephrase` `table` | 30M | 0.10 | ODC-BY | Tabularβnarrative realisation: reading quantities, units and row/column structure. | | |
| | 11 | `HuggingFaceCode/stack-v3-train`, `license_type=="permissive"`, per-language cap | 130M | 0.52 | ODC-BY + per-file SPDX | Code for compositional syntax and deterministic multi-step logic. **Must** filter to permissive licences (constraint 4) and cap XML/HTML/JSON/JS, which the card notes are byte-dominant. | | |
| | 12 | `SimpleStories/SimpleStories` | 40M | 0.07 | MIT | Syntactically simple long-range narrative. Disproportionately useful below ~200M params, where the model otherwise never sees a consistent referent across a full context window. | | |
| | 13 | `common-pile/libretexts` + `common-pile/arxiv_abstracts` | 40M | 0.08 | CC0 / per-doc PD/CC-BY | Openly-licensed OER textbooks and scientific abstracts; cleanest licence hygiene in the pool (arXiv metadata is CC0). | | |
| | | **total raw** | **1,310M** | **β4.35 GB** | | | | |
| **Headroom, and why it is sized like this.** Β§Phase 2 requires slack above the target for dedup loss | |
| and a held-out split. Cross-source near-dedup is expected to cost **12β18 %** (FinePhrase's four configs | |
| are rewrites of the *same* ~338.7M source documents, so they must be deduplicated against the shared | |
| `id`/`url` or they buy topic echo rather than knowledge; #2/#5/#6 overlap on encyclopedic topics; | |
| #1/#2 overlap on educational web content). 1,310M Γ 0.85 β 1,114M, minus a 2 % validation holdout | |
| (β22M) β **~1.09 B available**, which brackets the 1.0 B target and leaves room to land at 0.9 B | |
| without re-sourcing if Gate 3's throughput measurement forces the pre-registered fallback. | |
| **Sits comfortably inside the instance.** β4.35 GB of shards + ~2 GB of Arrow/tokeniser working | |
| overhead + 2 GB of final token store βͺ the measured **19.5 GB** `/kaggle/working`. Every source here is | |
| pullable as individual 270 MBβ3 GB shards, so the working directory is never the binding constraint β | |
| provided the build pulls shard-by-shard and deletes after tokenising, which Β§4 requires it to do. | |
| ## 3. Excluded, with reasons (so nobody re-litigates from scratch) | |
| | excluded | why | | |
| |---|---| | |
| | `allenai/c4`, `HuggingFaceFW/fineweb`, `Skylion007/openwebtext`, `tiiuae/falcon-refinedweb` | raw/unfiltered crawls β Β§3.5 | | |
| | `openai/gsm8k`, `cais/mmlu`, `allenai/ai2_arc`, `hellaswag`, `piqa`, `winogrande`, `truthfulqa` | the eval tasks themselves β Β§3.3. Also: never glob a directory that could catch these. | | |
| | `nvidia/Nemotron-CC*`, `-Math-v1`, `-Code-v1` | `gated: manual`; anonymous card and file reads return *"Access to dataset β¦ is restricted"* β token counts unverifiable and republication status unclear | | |
| | `openbmb/UltraData-Math` | card validates on the eval suite; exam-shaped L3 tiers; en+zh mixed | | |
| | `nvidia/Nemotron-Math-v2` | solution-formatted AoPS/SE traces, GSM8K-adjacent; reads as instruction data | | |
| | `allenai/peS2o` | ~11 GB per Btok by measurement β worst ratio in the pool for 42 B tokens we do not need | | |
| | `HuggingFaceTB/smollm-corpus` `cosmopedia-v2` | duplicate role with #3; also a documented card-vs-data mismatch (card 17.8 B vs ~31.7 B implied by stored `token_length`) | | |
| | `HuggingFaceFW/clean-wikipedia` | exists but README body is literally "Please see FineWiki instead" β deprecated | | |
| | `HuggingFaceTB/smollm-corpus` `python-edu` | **not a corpus.** Verified schema is `blob_id, repo_name, path, length_bytes, score, int_score` β a filter list pointing at gated `bigcode/the-stack-v2`, with no `text` column. Do not budget code tokens from it. | | |
| | `bigcode/the-stack-v2-dedup`, `codeparrot/github-code` | not verified this session; `the-stack-smol` is `gated: auto` | | |
| | `TinyLlama` 32k tokenizer | Llama-2-derived; redistribution licence status unverified β excluded from both tokenizer and any data | | |
| | `Cion-lab/Pretraining-10B-v1` (pre-existing in the account) | composition/mix unknown, token counts estimated with the **Qwen2.5** tokenizer, provenance not inspectable β cannot be cleared under Β§3.5 or Β§3.3 | | |
| ## 4. Card-vs-data warning | |
| Two sources disagree with themselves: `cosmopedia-v2` (17.8 B card vs ~31.7 B stored) and `finephrase` | |
| (486 B "completion tokens" card vs ~1.4 B implied by stored rows Γ inherited `token_count`). FinePhrase's | |
| `token_count`/`score` columns describe the **source document, not the rewrite**. **Therefore: no token | |
| number in this project is ever taken from a card.** Counts come from running our own tokenizer over the | |
| bytes, which is also what the Gate 2 manifest will record. | |
| ## 5. Build pipeline (Phase 2 design, on Kaggle CPU) | |
| 1. **Pull shard-by-shard, delete after consuming.** 73β89 MB/s measured (35β89 across runs), so 4.35 GB | |
| arrives in minutes; the constraint is *peak residency*, not bandwidth. Never materialise the mix. | |
| 2. **Normalise + filter.** Language filter on the row-level field where present; `int_scoreβ₯4` for | |
| fineweb-edu; `license_type=="permissive"` + per-language cap for stack-v3; drop `raw_mediawiki`. | |
| 3. **Dedup, memory-bounded.** Exact: 64-bit hash of the normalised document. Near: MinHash + LSH over | |
| 13-gram shingles (matching finemath's documented practice), **processed in bands against a | |
| disk-backed index**, because the box is a 30 GiB cgroup with **no swap** β a billion sketches in one | |
| dict is how a free-CPU job dies at hour six. | |
| 4. **Contamination audit (Β§3.3).** Mechanical 13-gram overlap of the surviving mix against benchmark | |
| **validation/dev** material only β including **MMLU's `dev` split specifically**, since those are the | |
| few-shot exemplars the harness will use, so overlap there is contamination by construction. A script | |
| computes and writes counts per source; **no item text is read, displayed or summarised by the agent.** | |
| Any source failing a stated threshold is dropped or cleaned; result written up in `docs/`. **Test | |
| splits are not opened until Phase 6.** | |
| 5. **Tokenise** with the frozen SmolLM2 BPE (49,152 β Β§3 of `01-plan`). Measured cost on this shape: | |
| 660,943 tok/s across 4 cores β **1.31 B tokens in ~0.55 h**. Not a bottleneck; process-shard rather | |
| than thread-shard, since rayon scaling saturated at 2.7Γ on 4 cores. | |
| 6. **Shard for exact positional resumption (Β§Phase 2 requirement).** Concatenate token ids into fixed-size | |
| **`uint16`** shards (`vocab 49,152 < 65,536`, so uint16 is exact and halves the bytes): **40 M tokens = | |
| 80 MB per shard**, ~28 shards, ~2 GB total. Then a resume position is literally | |
| `(shard_index, token_offset)` β one integer pair, memmap'd, so `skip_first_batches` on a replaying | |
| `Trainer` becomes an index advance rather than a decode. This is the property that makes Β§3.1's | |
| "same data, same position" checkable instead of hoped for. | |
| 7. **Hold out validation:** ~22 M tokens sampled **proportionally across sources**, never whole documents | |
| (whole-document holdout leaks topic statistics and makes PPL optimistic), written as its own shard so | |
| validation PPL is bit-reproducible at Gate 4. | |
| 8. **Publish** as public `ounce100m-mix-v1` with a manifest recording: per-source and total token counts, | |
| shard count and size, tokenizer identity (repo + revision + sha256 of `tokenizer.json`), per-shard | |
| sha256, the dedup parameters, the audit counts, and the cursor semantics of step 6. | |
| ## 6. Licensing the derived mix | |
| #5 and #6 are **CC-BY-SA-4.0** (share-alike), and `wikimedia` additionally carries GFDL. A published | |
| derivative containing them must be offered under SA with attribution and a change log. That is compatible | |
| with Β§2's "repos are public", so the plan is: **license the mix CC-BY-SA-4.0** and ship an | |
| `ATTRIBUTION.md` listing every source, its licence, and the transformations applied. ODC-BY sources | |
| (#1, #2, #7β#10) require notices kept and changes documented; `stack-v3-train` requires the per-file | |
| SPDX filter of Β§2 step 11 so we never republish a non-permissive file. | |
| **If a fully permissive mix is later required**, drop #5 and #6 (140 M raw) and backfill from #1/#3. The | |
| trade is a small loss of encyclopedic density for a licence with no reciprocal obligation β recorded | |
| here so the choice is available and deliberate rather than discovered under time pressure. | |
| ## 7. Gate 2 checklist, restated as things that must be true | |
| - [ ] dataset live on the Hub as a public `ounce100m-*` repo | |
| - [ ] token count **verified against the manifest** by an independent recount, not by the builder's own number | |
| - [ ] contamination audit passed **and written up** (per-source overlap counts, validation/dev only) | |
| - [ ] validation split reserved, its token count recorded | |
| - [ ] shard layout proven small enough that a training job reads a few shards at a time and never | |
| materialises the mix β demonstrated by a job that reports peak `/kaggle/working` usage while | |
| consuming shards, not asserted from the arithmetic above | |
| - [ ] exact-position cursor `(shard_index, token_offset)` demonstrated to be reproducible | |
| ## 8. Risks | |
| 1. **Dedup is the only stage with real wall-clock cost** and the only one that can exhaust RAM. Design | |
| for bands + disk; measure on one source before all thirteen. | |
| 2. **Unauthenticated Hub rate limits.** Jobs currently read anonymously and the API warns about it; at | |
| ~50 shard downloads that is probably fine, but 429s mid-build are plausible. Mitigation is the | |
| mounted-credential path (Β§8 of `01-plan`), which is being verified now. | |
| 3. **Card token counts are unreliable** (Β§4) β every number in this file is either measured over HTTP or | |
| will be replaced by our own tokeniser's count at Gate 2. | |
| 4. **`finepdfs-edu` adds Common Crawl's Terms of Use** on top of ODC-BY. Included because the subset is | |
| classifier-selected and globally deduplicated rather than a dump, but it is the least clean licence | |
| admission in the mix and the easiest to cut if the attribution review gets nervous. | |
| 5. **Two sources are rewrites/synthetic** (#3, #4, #9, #10). That is permitted explicitly by Β§3.5, but a | |
| 34 % synthetic share can produce fluent-but-homogeneous prose; the proportions above keep organic | |
| English (sources 1, 2, 5, 6, 8, 12, 13) in the majority, and that balance is a deliberate choice | |
| worth revisiting only with evidence. | |
| --- | |
| ## 9. Phase 2 step 1 ran β Β§2 corrected by measurement | |
| `dodosoomro/ounce100m-p2-source-inventory` (CPU, 116 s, 1,500 rows per source, tokenised with the | |
| **frozen SmolLM2 tokenizer**: vocab 49,152 confirmed, decode round-trips, 5.88 chars/token on a formal | |
| probe sentence). **13 of 18 entries resolved; 5 failed on config names, and two sources turned out to be | |
| unusable as described.** Β§2's estimates were right in aggregate (predicted β4.35 GB, measured 4.46 GB for | |
| the 570 M tokens that resolved) but wrong in specifics, so the specifics are replaced rather than | |
| averaged. | |
| ### 9.1 Config-name corrections (the datasets-server truth, from the error messages themselves) | |
| | Β§2 said | correct value | note | | |
| |---|---|---| | |
| | `fineweb-edu` `sample/10BT` | **`sample-10BT`** | hyphen, not a path. Full set: `default`, `sample-10BT`, `sample-100BT`, `sample-350BT`, plus per-dump `CC-MAIN-2025-05/08/13/18/21/26β¦` | | |
| | `finemath` `finemath4plus` | **`finemath-4plus`** | also `finemath-3plus`, `infiwebmath-3plus`, `infiwebmath-4plus` | | |
| | `finewiki` `data/enwiki` | config **`en`** | configs are plain language codes (`ab`, `ace`, `af`, β¦), not directory paths | | |
| | `wikipedia-monthly` `20250702.en` | **`20260101.en`** | the 2025-07 config is gone; **newer** dumps now exist (`20260101.*`), which is better for currency anyway | | |
| | `stack-v3-train` `python` | **`default`** only | **there are no per-language configs.** Language is a row field, and rows are whole repositories with a nested `files` struct | | |
| Every one of these would have been a mid-build failure. Recorded in `docs/00-platform-notes.md` style: | |
| the card's own naming is not an API β `load_dataset`'s error is. | |
| ### 9.2 Measured yield per document β this is what the token arithmetic actually runs on | |
| | source | tokens/doc | chars/doc | chars/token | non-English-flagged rows | verdict | | |
| |---|---|---|---|---|---| | |
| | `finepdfs-edu` `eng_Latn` | **1,321** | 14,352 | 3.98 | **128/1500 = 8.5 %** | keep, **must add our own language filter** | | |
| | `cosmopedia` `stanford` | 932 | 4,917 | 5.30 | 15/1500 = 1.0 % | keep | | |
| | `cosmopedia` `wikihow` | 907 | 4,541 | 4.97 | 0.7 % | keep | | |
| | `cosmopedia` `openstax` | 765 | 3,820 | 4.93 | 0.8 % | keep | | |
| | `cosmopedia` `auto_math_text` | 689 | 2,811 | 4.14 | 0.8 % | keep | | |
| | `cosmopedia` `khanacademy` | 677 | 3,047 | 4.26 | 1.1 % | keep | | |
| | `open-web-math` `default` | **1,182** | 6,795 | **3.31** | **112/1500 = 7.5 %** | keep, needs en filter; low chars/token = LaTeX density, which is the point | | |
| | `finephrase` `tutorial` | 779 | 4,903 | 4.58 | 1.1 % | keep. Column is **`text`**, not `completion` | | |
| | `finephrase` `faq` | 744 | 4,965 | 4.60 | 1.1 % | keep | | |
| | `finephrase` `table` | 715 | 4,772 | 4.59 | 1.1 % | keep | | |
| | `SimpleStories` | **285** | 1,260 | 4.33 | 0.5 % | keep, but it is 4Γ the *row* count for the same tokens | | |
| | `common-pile/arxiv_abstracts` | 185 | 791 | 4.31 | 2.2 % | keep. Column is **`text`**, not `raw_content` | | |
| | `common-pile/libretexts` | **75** | 964 | **2.42** | **1396/1500 = 93 %** | **DROP** | | |
| **`libretexts` is disqualified, and the reason matters.** Its `text` field is not prose: 93 % of sampled | |
| rows failed a naive English test, chars/token collapsed to 2.42, and the first example is a bare | |
| `https://math.libretexts.org/Bookshelves/...` URL. The Common-Pile schema puts real content in | |
| `metadata`/`raw_content` behind a per-source convention, so the column we assumed is an identifier field. | |
| Rather than reverse-engineer four sub-corpora's schemas for 10β20M tokens of marginal value, it is cut. | |
| `arxiv_abstracts` from the same publisher *did* work (its `text` is real abstract prose), so this is a | |
| per-dataset fact and not a Common-Pile-wide one. | |
| ### 9.3 Two open problems this created, and how they are being handled | |
| 1. **Code no longer has a cheap path.** With no per-language config, getting 60β130M tokens of Python/JS | |
| from `stack-v3-train` means streaming whole-repo rows, opening a nested `files` struct, and filtering | |
| on a row field β for a source that concentrates in XML/HTML/JSON anyway. Code is not mandatory (Β§3.4 | |
| makes proportions our call), and at 100M params / 1B tokens the marginal value is genuinely unclear. | |
| **Action: a dedicated mini-probe before committing** β measure rows-sampled-per-Python-repo and the | |
| nested schema β and hold code at **β€60M** in the plan until that returns. If it is awkward, drop it; | |
| that is a defensible Β§3.4 choice, not a failure. | |
| 2. **`eng_Latn` and `open-web-math` are not actually English-only** (8.5 % and 7.5 % of sampled rows fail | |
| a crude English test; the finepdfs sample included Cyrillic inline in a single "English" document). | |
| Β§3.4 is English-only, so **a language filter is now a required pipeline stage, not a courtesy** β | |
| applied per row on our side, on top of whatever the upstream config claims. `finepdfs` carries | |
| `full_doc_lid` / `per_page_languages` columns to do this cheaply; the general fallback is | |
| `language_score` where present and a fastText-style gate where not. | |
| ### 9.4 Revised totals | |
| Confirmed-and-usable now: 570M target tokens across 13 resolved configs, **β4.5 GB**, at measured yields | |
| of roughly **909k documents** β plus `fineweb-edu`/`finewiki`/`wikipedia-monthly`/`finemath` recoverable | |
| with the corrected config names (Β§9.1), which restores ~570M of the planned 1,310M. Net effect on the | |
| plan: **the shape survives, the line items do not.** Β§2's per-source token targets stay as targets; | |
| Β§2's config strings and two column names are superseded by this section. | |
| **Directly measured consequence for the build:** 909k documents for 570M tokens β 627 tokens/document | |
| average, so the mix is ~1.6β2.1M documents for 1.0β1.3B tokens. That is the number the dedup design has | |
| to be sized against β under 2.1M documents, a MinHash signature table comfortably fits the 30 GiB RAM, | |
| which materially de-risks the one stage Β§8 called the real cost. | |
| ## 10. Builder implemented, and one blocking flaw found and fixed | |
| `code/build/build_mix.py` (mirrored in `Cion-lab/ounce100m-code`) is the Phase 2 builder. It stages each | |
| source separately (stream β normalise β gates β dedup β tokenise) and then **merges** the staged shards | |
| into the final mixed shards by largest-deficit round-robin over documents. | |
| **E-016, and why the smoke test was mandatory.** The first version consumed one source to its target and | |
| only flushed a shard when the buffer filled. Since fineweb-edu's target alone is ~37 shards of 8M tokens, | |
| the first third of the mix would have been *nothing but web text*, then textbooks, then math β a curriculum | |
| drift across the run that no LR schedule can repair, and one the builder itself reported as success | |
| (checksums, headers, token counts and cursor probes were all clean). The verifier caught it as | |
| `distinct_sources_per_shard_min: 1`. The shuffle that was supposed to fix ordering only permuted documents | |
| *inside* a single-source buffer. After the two-stage rewrite, the same smoke run reports | |
| **`min = max = median = 4`** sources in every shard, 12,001,386 tokens recounted from the bytes, | |
| 0 out-of-range ids, 40/40 cursor probes and 40/40 document-boundary reconstructions correct. | |
| **Shard format** (shared by staged and final, so document boundaries survive to the trainer): | |
| ``` | |
| header : <IIII = vocab, n_docs, n_tokens, flags 16 B | |
| offsets : uint32 Γ (n_docs + 1), cumulative, [0] = 0 | |
| ids : uint16 Γ n_tokens (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes) | |
| doc i = ids[offsets[i] : offsets[i+1]] | |
| ``` | |
| Target 8M tokens/shard (16 MB, ~150 shards, ~2.5 GB for the mix). The resume cursor is | |
| `(shard_index, token_offset)` β one integer pair, which is what makes Β§3.1's "same data, same position" | |
| checkable and what keeps `Trainer`'s replaying `skip_first_batches` cheap on a map-style read. | |
| **Interruption safety, which was not optional.** `/kaggle/working` is destroyed when a session ends, and a | |
| full staging pass is hours of streaming. So each source is published to `Cion-lab/ounce100m-mix-stage` | |
| (the *stage* repo, distinct from the final `ounce100m-mix-v1`) the moment it completes, with a | |
| `record.json` holding its shard list, token count and drop statistics. A new session rebuilds its stage | |
| bookkeeping from the Hub and skips every source already published; only the source in flight at the moment | |
| of the kill is re-staged. `code/build/smoke_hub_resume.py` rehearses this by wiping the local tree between | |
| stage and merge, because "merge-only with an empty state and no restore path" is a failure mode that exits | |
| with code 0 while emitting an empty mix. | |
| **Revised raw total: 1.23 B tokens = 23 % headroom** over the 1.0 B trained target (raised fineweb-edu to | |
| 340M, finepdfs-edu to 180M, finemath to 160M, because Β§9's target edits had quietly pulled the plan from | |
| 1.31 B down to ~1.13 B β about 11 % headroom, which is not headroom once a 12β18 % dedup loss is applied). | |
| 16 sources, no code source yet (Β§9.3). | |
| **Gates the build must satisfy before Gate 2 is claimed**, as assertions in `verify_mix.py` rather than | |
| prose: every shard sha256 matches the manifest; no trailing bytes; offsets well-formed; zero out-of-range | |
| ids; total tokens recounted from bytes equals the manifest; and **β₯ half the distinct sources present in | |
| every shard**. The last one exists because the first build passed everything else. | |
| ## 11. Interruption safety proven, and the build that this authorises | |
| A CPU session's `/kaggle/working` is destroyed when it ends. Everything below was measured, not assumed | |
| (`dodosoomro/ounce100m-p2-cold-resume-rehearsal` **v6**, 17:35Z, 55 s wall-clock, pinned REV `428c5234`). | |
| | Step | Observed | | |
| |---|---| | |
| | 1. stage two sources, publish each | 2,400,867 tokens staged; `record.json` + `shard-0000.bin` per source on the Hub | | |
| | 2. **wipe the local tree** (`root_exists: false`, `stage_dir_exists: false`) | faithful simulation of a session kill, not a soft restart | | |
| | 3. `--merge-only` with the Hub as the only state | `hub: restored 2 staged source(s) (2,400,867 tokens)` β merged shard 2,352,747 tok / 5,815 docs, **2/2 sources in it**, val shard 48,120 tok | | |
| | 4. `verify_mix.py` over the restored mix | `PASS: true`, `sha_ok 1/1`, `out_of_range_ids: 0`, `trailing_bytes: 0`, cursor probes **40/40 boundaries reconstructed, 0 mismatches**, `recount_matches_manifest: true` | | |
| | 5. verdict + self-cleanup | `COLD_RESUME_PASSED=True`, `cleanup=deleted`, throwaway repo confirmed gone from the namespace | | |
| v4 of the same rehearsal had reported `false` over an identically-working pipeline: a stray | |
| `print("shards=", β¦)` inside the `BUILD_JSON` block made the parent's `json.loads` fail and a helper fold | |
| that into `None` (memory/ERRORS.md **E-019**, and the report-shape recurrence **E-023**). The fix was to | |
| the harness, plus a `(obj, why_null)` return and a `manifest.json` fallback so a verdict can never again | |
| be silently null. | |
| **The build now running** (`dodosoomro/ounce100m-p2-full-mix-build`, 17:38Z, CPU, `sessionTimeoutSeconds | |
| 43200`, pinned REV `98e49560`): `build_mix.py --root /kaggle/working/mixroot --hub-repo | |
| Cion-lab/ounce100m-mix-stage` with defaults β stage cap 1.31 B tokens, per-source targets summing to | |
| 1.23 B, 8 M tokens/mixed shard, `VAL_STRIDE 50` holdout, shuffle seed 20260919. Each source publishes as | |
| it completes, so an interruption costs at most one source's streaming time, and a re-launch restores the | |
| finished ones from the Hub (step 3 above is that exact path). | |
| **Audit semantics, hardened the same way.** `audit_contamination.py` compares 13-token windows against | |
| `{train, validation, dev}` of the eight tasks β never `test` β and emits counts only. It now also reports | |
| `shards_missing`, and its verdict is composite: `AUDIT_PASSED` requires every manifest shard present and | |
| read, > 0 documents and grams measured, β₯ 6 readable reference sets, zero reference errors, and zero | |
| overlap. Exit 6 means *overlap found* (a real finding); exit 7 means *the audit did not measure* (a bug). | |
| A zero-overlap report over shards that were not on disk is no longer expressible. | |
| **Publishing ships the tokenizer.** `publish_mix.py` copies `tokenizer.json` into the dataset root, | |
| re-hashes it and aborts with exit 8 unless it matches the sha256 recorded in the manifest β so the | |
| published ids and the published tokenizer are provably the pair that produced them (Β§3.12). | |
| ## 12. Gate 2 β the record, including where it failed and what that failure was worth | |
| **Merge and verify: passed, twice, reproducibly.** From the 15 staged sources | |
| (`Cion-lab/ounce100m-mix-stage`, 1,150,015,841 tokens staged): | |
| | field | value | | |
| |---|---| | |
| | merged shards | **141** (each drawn from all 15 sources) | | |
| | train tokens | **1,127,163,952** | | |
| | held-out validation | **22,851,889** in 12 shards (every 50th merged document) | | |
| | documents | 1,078,807 | | |
| | per-source share | 75,144,263 tokens each, equal by construction | | |
| | merge time | 16.4 s (restore from the Hub included: 123.8 s) | | |
| | `verify_mix` | **`PASS: true`** β sha256 141/141, headers 141/141, 0 bad offsets, 0 out-of-range ids, 0 trailing bytes, recount matches manifest, 40/40 cursor probes, `distinct_sources_per_shard_min = 15` (floor was 7) | | |
| | `finewiki__en` | **never staged.** Gate 2's own diagnostic showed the columns present and `first_row_ok: true` with a truncated `ImportError` in the loading script. 15 sources at 1.127 B is above the 1.0 B target with 12.7 % headroom, so this is a composition note, not a blocker | | |
| **Then the audit refused to publish, and it was right to.** `audit_contamination` exited 7: | |
| | task | 13-gram hits | reference size (verified against the dataset server) | | |
| |---|---|---| | |
| | mmlu | 904,288 | dev 285 + validation 1,531 items β 149k grams | | |
| | gsm8k | 185,600 | train 7,473 items = 360,656 grams | | |
| | hellaswag | 1,852 | train 39,905 + validation 10,042 | | |
| | arc_easy | 358 | train 2,251 + validation 570 | | |
| | arc_challenge | 214 | train 1,119 + validation 299 | | |
| | piqa | **not measured** | the loader yielded 0 of 16,113 train rows without raising | | |
| | winogrande | (not in the visible excerpt) | train 40,398 + validation 1,267 | | |
| | truthfulqa | (repo name was wrong) | validation 817, `multiple_choice` | | |
| `expected_false_positives` for the whole scan was **β0.0001**, so these are matches, not collisions. The | |
| first reading I wrote in `STATE.md` β "GSM8K's whole train split is in the mix" β was an over-claim: hits | |
| are counted per *mix* gram, so a small number of frequently repeated reference grams can produce a large | |
| count, and 185,600 hits against 360,656 distinct GSM8K grams means "up to half of its gram vocabulary | |
| occurs somewhere in 1.1 B tokens". The columns that settle severity are `mix_documents_majority_hit` | |
| (documents whose grams are more than half matched β verbatim inclusion) and `source_hit_document_fraction` | |
| (which source, and how much of it) β and both are only computable on the fixed audit, which is what Gate | |
| 2c/2d exist to produce. | |
| **Why upstream did not catch it, verified rather than assumed:** FineMath's card states it removed 13-gram | |
| overlap against the **test** sets of GSM8k, MATH, MMLU and ARC; Cosmopedia's describes a 10-gram + | |
| `difflib` pass against **test benchmarks**; OpenWebMath documents only SimHash self-deduplication with no | |
| benchmark claim. `validation`/`dev`/`train` overlap was therefore never in scope for any of them β and | |
| `validation` is the split lm-eval scores HellaSwag, PIQA, Winogrande and TruthfulQA on | |
| (`docs/05-eval-plan.md` Β§7). | |
| **Gate 2 remained unsigned at the end of this stage.** `Cion-lab/ounce100m-mix-v1` did not exist yet. The | |
| path to it: attribution β | |
| drop the offending documents by mask (or a source if attribution says it is broadly soaked) β re-merge | |
| with `--exclude-dir` β re-audit until `AUDIT_PASSED` with all 8 tasks covered β `publish_mix`, which now | |
| refuses to ship unless the audit passed, covered all 8, scanned the val shards, and its `mix_bytes_sha256` | |
| matches the exact manifest being published. `E-030` explains why the first filtered re-merge could not | |
| work and what was changed. | |
| ### 12.4 Resolution β the filtered mix, measured (Gates 2bβ2f, 2026-09-19) | |
| **Attribution first, always.** Gate 2b/2c scanned the *staged* shards per source, which is the only way to | |
| say where contamination lives rather than how much the merged mix contains. Of 1,955,832 staged documents, | |
| **10,425 (0.533 %)** carried at least one matched 13-gram and **0 were majority-matched** β shared spans, | |
| not wholesale inclusion. The count concentrated in five sources: finemath-4plus 2,688 documents, fineweb-edu | |
| 1,043, cosmopedia `auto_math_text` 568, finepdfs-edu 420, open-web-math 339, with the remaining ~5,367 | |
| spread over ten more. Per-source *hit counts* summed to ~50,700 while the merged scan reported 2,096,540 for | |
| essentially the same gram volume; the discrepancy is unreconciled and recorded as **E-032c** rather than | |
| smoothed over. It changes no decision: the exclusion is per-document and the post-filter overlap is zero | |
| either way. | |
| **The exclusion, and what it cost.** `audit_contamination --write-filter` emitted one `.u8` mask per source | |
| (1 byte per staged document, ~700 KB total) marking documents that contain a matched window. Re-merging with | |
| `--exclude-dir` dropped **5,878 documents = 17,366,967 tokens** (1.5 % of the mix) and the re-audit of the | |
| result gave **`overlap_total = 0`** with 8/8 benchmark tasks covered, the held-out validation shards | |
| scanned clean, and `MEASURED: true`. Final mix: **1,109,714,831 training tokens** in 141 shards plus | |
| 22,851,889 held-out validation tokens β still **11.0 %** above the 1.0 B target, so the pre-registered | |
| 0.9 B fallback (D-012) did not fire and the contamination fix cost nothing that the token budget had to | |
| pay for. | |
| **Reference coverage is now a precondition, not an assumption.** Gate 2b found that `piqa`'s train and | |
| validation splits returned **0 rows** from the `datasets` library inside the Kaggle image while the server | |
| reports 16,113 + 1,838 β the first audit had scored that as "clean". The audit now fetches each split's | |
| server-side row count (`/size`), refuses any split it read to under 90 % of it, retries through the HTTP | |
| `/rows` endpoint, and marks the whole report `MEASURED: false` if a loader still falls short. TruthfulQA's | |
| answer text also sits under keys (`mc1_targets`, `mc2_targets`, `abstractions`) the first version never | |
| read. All reference splits used are train/validation/dev; **no test split was opened** (Β§3.3). | |
| **Why publication took three attempts (E-033).** Gates 2e and 2f both built and audited the mix correctly | |
| and then died in the uploader: `publish_mix` sent one commit per file, and the Hub began rejecting every | |
| request β including 30-byte files β after ~139 commits. That is a commit-rate ceiling, not a network fault, | |
| and the "resumable" re-send was itself broken because `RepoFile.lfs` is a dict and `getattr(ent.lfs, | |
| "sha256", None)` is therefore `None` for every shard, i.e. *nothing* looked published. The uploader now | |
| reads `lfs["sha256"] or lfs["oid"]`, sends the remaining files as **one `upload_folder` commit**, and after | |
| the commit re-lists the repo and re-hashes, exiting 11 if a path is absent and 12 if a remote digest | |
| disagrees with the local bytes. `Cion-lab/ounce100m-mix-v1` held 142 of 173 files when Gate 2g re-ran the | |
| same pipeline with it. | |
| **Gate 2 is signed** (`dodosoomro/ounce100m-gate-2g-publish-mix-v1`, complete 2026-09-20T00:31Z, CPU, zero | |
| GPU quota). The rebuilt mix reproduced Gate 2e's token total exactly β **1,109,714,831** β which is the | |
| determinism claim in Β§11 tested a second time on a different session with the same inputs. `audit.json` on | |
| the Hub reports `overlap_total = 0`, `mix_documents_with_any_hit = 0`, `majority_hit_docs = 0`, | |
| `expected_false_positives = 0.0`, `tasks_covered` = all eight, `val_hits = 0` across 21,898 held-out | |
| documents, and `AUDIT_PASSED`, `MEASURED`, `COVERED_ALL_SHARDS` all true, bound to these exact bytes by | |
| `mix_bytes_sha256 = 732eec7f19e77f36β¦`. `verify_mix` exited 0. The batched uploader landed all 174 files | |
| (139 train shards, 12 val shards, 16 `filter/` objects, `manifest.json`, `audit.json`, `tokenizer.json`, | |
| `README.md`, `ATTRIBUTION.md`, `build_state.json`, `.gitattributes`), and four train plus two validation | |
| shards were then read back from the anonymous `resolve` endpoint and matched the manifest's sha256 β so the | |
| mix is public and serving, not merely listed (Β§3.13). | |
| One editorial note on this section: Β§12.2's first draft treated `overlap_total` as a count of contaminated | |
| documents. It is a count of matched *grams*, which is why a 41Γ difference between the per-source and | |
| merged hit totals could coexist with a 1.3 % agreement on document counts (E-032c). The action taken β | |
| drop the documents, re-audit to zero β was unaffected, but the number in the record was read wrongly for | |
| two gates, and the fix was to write the density columns (`mix_documents_with_any_hit`, | |
| `mix_documents_majority_hit`) and quote *those* when judging severity. | |