# 02 — Mix plan: sources, counts, rationale, and the build Phase 1 deliverable. Target: **1.0 B trained tokens** (§2 band 900M–1.1B; frozen number confirmed at Gate 3). Every repo id, size, date and licence below was read live on 2026-09-19 from the Hub card, the `/api/datasets/` metadata endpoint, or parquet footers over HTTP — not from memory. Where a card and the stored data disagree, that is recorded rather than smoothed over. ## 1. Selection rules applied, operationally - **§3.4 English only.** Every source is either English-config-scoped (`20250702.en`, `eng_Latn`, `sample/10BT`) or carries a per-row language field we filter on. Code and math are allowed and are deliberately included. - **§3.5 No raw crawls.** Excluded as a class: `allenai/c4`, `HuggingFaceFW/fineweb`, `Skylion007/openwebtext`, `tiiuae/falcon-refinedweb`. What is admitted instead is **quality-scored** (`fineweb-edu` carries a per-row classifier `score`/`int_score`), **curated/rewritten** (`finewiki`, `cosmopedia`), **synthetic-textbook** (`finephrase`), or **openly-licensed domain corpora** (`common-pile/*`, `stack-v3-train`). Common-Crawl appears only as the *provenance* of a classifier-filtered subset, never as the selection criterion. - **§3.3 No contamination.** Two mechanisms, not one assertion: 1. **Prefer sources with published decontamination.** `HuggingFaceTB/finemath` documents 13-gram removal against GSM8K, MATH, MMLU and ARC test sets, with a public audit log (`HuggingFaceTB/finemath_contamination_report`). 2. **Reject sources whose own cards validate on the eval suite.** That is why `openbmb/UltraData-Math` is **dropped entirely** despite being attractive on paper (apache-2.0, 170B+ tokens): its card lists MMLU/ARC-E/ARC-C/HellaSwag/PIQA/Winogrande/GSM8K/MATH500 as validation targets and its L3 tiers are synthetic *exam-shaped* text. Its en+zh mixing would also need a language filter. Similarly excluded: `nvidia/Nemotron-Math-v2` (AoPS/SE solution traces, i.e. adjacent to GSM8K's own source pool, and solution-formatted in a way that quietly steers toward instruction data). 3. The mechanical overlap audit of §5 runs against **validation/dev material only**; test splits stay untouched until Phase 6. ## 2. The mix: 1.31 B raw → ~1.05 B after dedup → ~1.0 B trained | # | Source (config) | raw tokens | GB to pull | licence | rationale, from corpus properties | |---|---|---|---|---|---| | 1 | `HuggingFaceFW/fineweb-edu` `sample/10BT`, keep `int_score≥4` | 300M | 1.1 | ODC-BY | Classifier-scored general English; the only cheap source of broad register coverage. Score-gating is a *quality* filter, which §3.5 permits, and it is the source's own documented axis. | | 2 | `HuggingFaceFW/finepdfs-edu` `data/eng_Latn` | 150M | 0.63 | ODC-BY + CC ToU | Globally-deduplicated long-form educational PDFs. Long document structure teaches coherence over >1k tokens, which chunked web prose does not. `eng_Latn` only (69 scripts exist). | | 3 | `HuggingFaceTB/cosmopedia` stanford+openstax+khanacademy+auto_math_text | 130M | 0.33 | **Apache-2.0** | Synthetic textbooks: causal and explanatory chains ("because", "therefore", worked steps). This is the register that builds science-QA and multi-step reasoning ability. Per-config GB measured, not estimated. | | 4 | `HuggingFaceTB/cosmopedia` `wikihow` | 30M | 0.05 | Apache-2.0 | Procedural how-to steps → ordered physical-world inference. Small slice, specific job. | | 5 | `HuggingFaceFW/finewiki` `data/enwiki` | 80M | 0.43 | **CC-BY-SA-4.0** | Link-resolved, rewritten encyclopedic prose; definitional and comparative sentence forms that factual/recall-style prompts reward. Heaviest GB/token in the mix, hence capped. | | 6 | `omarkamali/wikipedia-monthly` `20250702.en` | 60M | 0.03 | CC-BY-SA-4.0 | Cheapest encyclopedic tokens available (0.5 GB/Btok) and *current*; drop the `raw_mediawiki` column or it dominates the byte budget. | | 7 | `HuggingFaceTB/finemath` `finemath-4plus` | 130M | 0.25 | ODC-BY | Best math density per GB (1.9) **and** the only candidate with a published eval-set decontamination report. | | 8 | `open-web-math/open-web-math` | 70M | 0.15 | ODC-BY (card body) | Forum/webbook math in natural prose, incl. LaTeX. Complements #7's textbook skew with multi-turn human reasoning text. Oldest source here (2023-10) — cited as such. | | 9 | `HuggingFaceFW/finephrase` `tutorial` + `faq` | 90M | 0.31 | ODC-BY | Rewritten into stepwise-tutorial and question/answer registers — forms largely absent from organic prose. | | 10 | `HuggingFaceFW/finephrase` `table` | 30M | 0.10 | ODC-BY | Tabular→narrative realisation: reading quantities, units and row/column structure. | | 11 | `HuggingFaceCode/stack-v3-train`, `license_type=="permissive"`, per-language cap | 130M | 0.52 | ODC-BY + per-file SPDX | Code for compositional syntax and deterministic multi-step logic. **Must** filter to permissive licences (constraint 4) and cap XML/HTML/JSON/JS, which the card notes are byte-dominant. | | 12 | `SimpleStories/SimpleStories` | 40M | 0.07 | MIT | Syntactically simple long-range narrative. Disproportionately useful below ~200M params, where the model otherwise never sees a consistent referent across a full context window. | | 13 | `common-pile/libretexts` + `common-pile/arxiv_abstracts` | 40M | 0.08 | CC0 / per-doc PD/CC-BY | Openly-licensed OER textbooks and scientific abstracts; cleanest licence hygiene in the pool (arXiv metadata is CC0). | | | **total raw** | **1,310M** | **≈4.35 GB** | | | **Headroom, and why it is sized like this.** §Phase 2 requires slack above the target for dedup loss and a held-out split. Cross-source near-dedup is expected to cost **12–18 %** (FinePhrase's four configs are rewrites of the *same* ~338.7M source documents, so they must be deduplicated against the shared `id`/`url` or they buy topic echo rather than knowledge; #2/#5/#6 overlap on encyclopedic topics; #1/#2 overlap on educational web content). 1,310M × 0.85 ≈ 1,114M, minus a 2 % validation holdout (≈22M) → **~1.09 B available**, which brackets the 1.0 B target and leaves room to land at 0.9 B without re-sourcing if Gate 3's throughput measurement forces the pre-registered fallback. **Sits comfortably inside the instance.** ≈4.35 GB of shards + ~2 GB of Arrow/tokeniser working overhead + 2 GB of final token store ≪ the measured **19.5 GB** `/kaggle/working`. Every source here is pullable as individual 270 MB–3 GB shards, so the working directory is never the binding constraint — provided the build pulls shard-by-shard and deletes after tokenising, which §4 requires it to do. ## 3. Excluded, with reasons (so nobody re-litigates from scratch) | excluded | why | |---|---| | `allenai/c4`, `HuggingFaceFW/fineweb`, `Skylion007/openwebtext`, `tiiuae/falcon-refinedweb` | raw/unfiltered crawls → §3.5 | | `openai/gsm8k`, `cais/mmlu`, `allenai/ai2_arc`, `hellaswag`, `piqa`, `winogrande`, `truthfulqa` | the eval tasks themselves → §3.3. Also: never glob a directory that could catch these. | | `nvidia/Nemotron-CC*`, `-Math-v1`, `-Code-v1` | `gated: manual`; anonymous card and file reads return *"Access to dataset … is restricted"* → token counts unverifiable and republication status unclear | | `openbmb/UltraData-Math` | card validates on the eval suite; exam-shaped L3 tiers; en+zh mixed | | `nvidia/Nemotron-Math-v2` | solution-formatted AoPS/SE traces, GSM8K-adjacent; reads as instruction data | | `allenai/peS2o` | ~11 GB per Btok by measurement — worst ratio in the pool for 42 B tokens we do not need | | `HuggingFaceTB/smollm-corpus` `cosmopedia-v2` | duplicate role with #3; also a documented card-vs-data mismatch (card 17.8 B vs ~31.7 B implied by stored `token_length`) | | `HuggingFaceFW/clean-wikipedia` | exists but README body is literally "Please see FineWiki instead" → deprecated | | `HuggingFaceTB/smollm-corpus` `python-edu` | **not a corpus.** Verified schema is `blob_id, repo_name, path, length_bytes, score, int_score` — a filter list pointing at gated `bigcode/the-stack-v2`, with no `text` column. Do not budget code tokens from it. | | `bigcode/the-stack-v2-dedup`, `codeparrot/github-code` | not verified this session; `the-stack-smol` is `gated: auto` | | `TinyLlama` 32k tokenizer | Llama-2-derived; redistribution licence status unverified → excluded from both tokenizer and any data | | `Cion-lab/Pretraining-10B-v1` (pre-existing in the account) | composition/mix unknown, token counts estimated with the **Qwen2.5** tokenizer, provenance not inspectable → cannot be cleared under §3.5 or §3.3 | ## 4. Card-vs-data warning Two sources disagree with themselves: `cosmopedia-v2` (17.8 B card vs ~31.7 B stored) and `finephrase` (486 B "completion tokens" card vs ~1.4 B implied by stored rows × inherited `token_count`). FinePhrase's `token_count`/`score` columns describe the **source document, not the rewrite**. **Therefore: no token number in this project is ever taken from a card.** Counts come from running our own tokenizer over the bytes, which is also what the Gate 2 manifest will record. ## 5. Build pipeline (Phase 2 design, on Kaggle CPU) 1. **Pull shard-by-shard, delete after consuming.** 73–89 MB/s measured (35–89 across runs), so 4.35 GB arrives in minutes; the constraint is *peak residency*, not bandwidth. Never materialise the mix. 2. **Normalise + filter.** Language filter on the row-level field where present; `int_score≥4` for fineweb-edu; `license_type=="permissive"` + per-language cap for stack-v3; drop `raw_mediawiki`. 3. **Dedup, memory-bounded.** Exact: 64-bit hash of the normalised document. Near: MinHash + LSH over 13-gram shingles (matching finemath's documented practice), **processed in bands against a disk-backed index**, because the box is a 30 GiB cgroup with **no swap** — a billion sketches in one dict is how a free-CPU job dies at hour six. 4. **Contamination audit (§3.3).** Mechanical 13-gram overlap of the surviving mix against benchmark **validation/dev** material only — including **MMLU's `dev` split specifically**, since those are the few-shot exemplars the harness will use, so overlap there is contamination by construction. A script computes and writes counts per source; **no item text is read, displayed or summarised by the agent.** Any source failing a stated threshold is dropped or cleaned; result written up in `docs/`. **Test splits are not opened until Phase 6.** 5. **Tokenise** with the frozen SmolLM2 BPE (49,152 — §3 of `01-plan`). Measured cost on this shape: 660,943 tok/s across 4 cores → **1.31 B tokens in ~0.55 h**. Not a bottleneck; process-shard rather than thread-shard, since rayon scaling saturated at 2.7× on 4 cores. 6. **Shard for exact positional resumption (§Phase 2 requirement).** Concatenate token ids into fixed-size **`uint16`** shards (`vocab 49,152 < 65,536`, so uint16 is exact and halves the bytes): **40 M tokens = 80 MB per shard**, ~28 shards, ~2 GB total. Then a resume position is literally `(shard_index, token_offset)` — one integer pair, memmap'd, so `skip_first_batches` on a replaying `Trainer` becomes an index advance rather than a decode. This is the property that makes §3.1's "same data, same position" checkable instead of hoped for. 7. **Hold out validation:** ~22 M tokens sampled **proportionally across sources**, never whole documents (whole-document holdout leaks topic statistics and makes PPL optimistic), written as its own shard so validation PPL is bit-reproducible at Gate 4. 8. **Publish** as public `ounce100m-mix-v1` with a manifest recording: per-source and total token counts, shard count and size, tokenizer identity (repo + revision + sha256 of `tokenizer.json`), per-shard sha256, the dedup parameters, the audit counts, and the cursor semantics of step 6. ## 6. Licensing the derived mix #5 and #6 are **CC-BY-SA-4.0** (share-alike), and `wikimedia` additionally carries GFDL. A published derivative containing them must be offered under SA with attribution and a change log. That is compatible with §2's "repos are public", so the plan is: **license the mix CC-BY-SA-4.0** and ship an `ATTRIBUTION.md` listing every source, its licence, and the transformations applied. ODC-BY sources (#1, #2, #7–#10) require notices kept and changes documented; `stack-v3-train` requires the per-file SPDX filter of §2 step 11 so we never republish a non-permissive file. **If a fully permissive mix is later required**, drop #5 and #6 (140 M raw) and backfill from #1/#3. The trade is a small loss of encyclopedic density for a licence with no reciprocal obligation — recorded here so the choice is available and deliberate rather than discovered under time pressure. ## 7. Gate 2 checklist, restated as things that must be true - [ ] dataset live on the Hub as a public `ounce100m-*` repo - [ ] token count **verified against the manifest** by an independent recount, not by the builder's own number - [ ] contamination audit passed **and written up** (per-source overlap counts, validation/dev only) - [ ] validation split reserved, its token count recorded - [ ] shard layout proven small enough that a training job reads a few shards at a time and never materialises the mix — demonstrated by a job that reports peak `/kaggle/working` usage while consuming shards, not asserted from the arithmetic above - [ ] exact-position cursor `(shard_index, token_offset)` demonstrated to be reproducible ## 8. Risks 1. **Dedup is the only stage with real wall-clock cost** and the only one that can exhaust RAM. Design for bands + disk; measure on one source before all thirteen. 2. **Unauthenticated Hub rate limits.** Jobs currently read anonymously and the API warns about it; at ~50 shard downloads that is probably fine, but 429s mid-build are plausible. Mitigation is the mounted-credential path (§8 of `01-plan`), which is being verified now. 3. **Card token counts are unreliable** (§4) — every number in this file is either measured over HTTP or will be replaced by our own tokeniser's count at Gate 2. 4. **`finepdfs-edu` adds Common Crawl's Terms of Use** on top of ODC-BY. Included because the subset is classifier-selected and globally deduplicated rather than a dump, but it is the least clean licence admission in the mix and the easiest to cut if the attribution review gets nervous. 5. **Two sources are rewrites/synthetic** (#3, #4, #9, #10). That is permitted explicitly by §3.5, but a 34 % synthetic share can produce fluent-but-homogeneous prose; the proportions above keep organic English (sources 1, 2, 5, 6, 8, 12, 13) in the majority, and that balance is a deliberate choice worth revisiting only with evidence. --- ## 9. Phase 2 step 1 ran — §2 corrected by measurement `dodosoomro/ounce100m-p2-source-inventory` (CPU, 116 s, 1,500 rows per source, tokenised with the **frozen SmolLM2 tokenizer**: vocab 49,152 confirmed, decode round-trips, 5.88 chars/token on a formal probe sentence). **13 of 18 entries resolved; 5 failed on config names, and two sources turned out to be unusable as described.** §2's estimates were right in aggregate (predicted ≈4.35 GB, measured 4.46 GB for the 570 M tokens that resolved) but wrong in specifics, so the specifics are replaced rather than averaged. ### 9.1 Config-name corrections (the datasets-server truth, from the error messages themselves) | §2 said | correct value | note | |---|---|---| | `fineweb-edu` `sample/10BT` | **`sample-10BT`** | hyphen, not a path. Full set: `default`, `sample-10BT`, `sample-100BT`, `sample-350BT`, plus per-dump `CC-MAIN-2025-05/08/13/18/21/26…` | | `finemath` `finemath4plus` | **`finemath-4plus`** | also `finemath-3plus`, `infiwebmath-3plus`, `infiwebmath-4plus` | | `finewiki` `data/enwiki` | config **`en`** | configs are plain language codes (`ab`, `ace`, `af`, …), not directory paths | | `wikipedia-monthly` `20250702.en` | **`20260101.en`** | the 2025-07 config is gone; **newer** dumps now exist (`20260101.*`), which is better for currency anyway | | `stack-v3-train` `python` | **`default`** only | **there are no per-language configs.** Language is a row field, and rows are whole repositories with a nested `files` struct | Every one of these would have been a mid-build failure. Recorded in `docs/00-platform-notes.md` style: the card's own naming is not an API — `load_dataset`'s error is. ### 9.2 Measured yield per document — this is what the token arithmetic actually runs on | source | tokens/doc | chars/doc | chars/token | non-English-flagged rows | verdict | |---|---|---|---|---|---| | `finepdfs-edu` `eng_Latn` | **1,321** | 14,352 | 3.98 | **128/1500 = 8.5 %** | keep, **must add our own language filter** | | `cosmopedia` `stanford` | 932 | 4,917 | 5.30 | 15/1500 = 1.0 % | keep | | `cosmopedia` `wikihow` | 907 | 4,541 | 4.97 | 0.7 % | keep | | `cosmopedia` `openstax` | 765 | 3,820 | 4.93 | 0.8 % | keep | | `cosmopedia` `auto_math_text` | 689 | 2,811 | 4.14 | 0.8 % | keep | | `cosmopedia` `khanacademy` | 677 | 3,047 | 4.26 | 1.1 % | keep | | `open-web-math` `default` | **1,182** | 6,795 | **3.31** | **112/1500 = 7.5 %** | keep, needs en filter; low chars/token = LaTeX density, which is the point | | `finephrase` `tutorial` | 779 | 4,903 | 4.58 | 1.1 % | keep. Column is **`text`**, not `completion` | | `finephrase` `faq` | 744 | 4,965 | 4.60 | 1.1 % | keep | | `finephrase` `table` | 715 | 4,772 | 4.59 | 1.1 % | keep | | `SimpleStories` | **285** | 1,260 | 4.33 | 0.5 % | keep, but it is 4× the *row* count for the same tokens | | `common-pile/arxiv_abstracts` | 185 | 791 | 4.31 | 2.2 % | keep. Column is **`text`**, not `raw_content` | | `common-pile/libretexts` | **75** | 964 | **2.42** | **1396/1500 = 93 %** | **DROP** | **`libretexts` is disqualified, and the reason matters.** Its `text` field is not prose: 93 % of sampled rows failed a naive English test, chars/token collapsed to 2.42, and the first example is a bare `https://math.libretexts.org/Bookshelves/...` URL. The Common-Pile schema puts real content in `metadata`/`raw_content` behind a per-source convention, so the column we assumed is an identifier field. Rather than reverse-engineer four sub-corpora's schemas for 10–20M tokens of marginal value, it is cut. `arxiv_abstracts` from the same publisher *did* work (its `text` is real abstract prose), so this is a per-dataset fact and not a Common-Pile-wide one. ### 9.3 Two open problems this created, and how they are being handled 1. **Code no longer has a cheap path.** With no per-language config, getting 60–130M tokens of Python/JS from `stack-v3-train` means streaming whole-repo rows, opening a nested `files` struct, and filtering on a row field — for a source that concentrates in XML/HTML/JSON anyway. Code is not mandatory (§3.4 makes proportions our call), and at 100M params / 1B tokens the marginal value is genuinely unclear. **Action: a dedicated mini-probe before committing** — measure rows-sampled-per-Python-repo and the nested schema — and hold code at **≤60M** in the plan until that returns. If it is awkward, drop it; that is a defensible §3.4 choice, not a failure. 2. **`eng_Latn` and `open-web-math` are not actually English-only** (8.5 % and 7.5 % of sampled rows fail a crude English test; the finepdfs sample included Cyrillic inline in a single "English" document). §3.4 is English-only, so **a language filter is now a required pipeline stage, not a courtesy** — applied per row on our side, on top of whatever the upstream config claims. `finepdfs` carries `full_doc_lid` / `per_page_languages` columns to do this cheaply; the general fallback is `language_score` where present and a fastText-style gate where not. ### 9.4 Revised totals Confirmed-and-usable now: 570M target tokens across 13 resolved configs, **≈4.5 GB**, at measured yields of roughly **909k documents** — plus `fineweb-edu`/`finewiki`/`wikipedia-monthly`/`finemath` recoverable with the corrected config names (§9.1), which restores ~570M of the planned 1,310M. Net effect on the plan: **the shape survives, the line items do not.** §2's per-source token targets stay as targets; §2's config strings and two column names are superseded by this section. **Directly measured consequence for the build:** 909k documents for 570M tokens ≈ 627 tokens/document average, so the mix is ~1.6–2.1M documents for 1.0–1.3B tokens. That is the number the dedup design has to be sized against — under 2.1M documents, a MinHash signature table comfortably fits the 30 GiB RAM, which materially de-risks the one stage §8 called the real cost. ## 10. Builder implemented, and one blocking flaw found and fixed `code/build/build_mix.py` (mirrored in `Cion-lab/ounce100m-code`) is the Phase 2 builder. It stages each source separately (stream → normalise → gates → dedup → tokenise) and then **merges** the staged shards into the final mixed shards by largest-deficit round-robin over documents. **E-016, and why the smoke test was mandatory.** The first version consumed one source to its target and only flushed a shard when the buffer filled. Since fineweb-edu's target alone is ~37 shards of 8M tokens, the first third of the mix would have been *nothing but web text*, then textbooks, then math — a curriculum drift across the run that no LR schedule can repair, and one the builder itself reported as success (checksums, headers, token counts and cursor probes were all clean). The verifier caught it as `distinct_sources_per_shard_min: 1`. The shuffle that was supposed to fix ordering only permuted documents *inside* a single-source buffer. After the two-stage rewrite, the same smoke run reports **`min = max = median = 4`** sources in every shard, 12,001,386 tokens recounted from the bytes, 0 out-of-range ids, 40/40 cursor probes and 40/40 document-boundary reconstructions correct. **Shard format** (shared by staged and final, so document boundaries survive to the trainer): ``` header :