ounce100m-code / docs /02-mix-plan.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
36 kB

02 β€” Mix plan: sources, counts, rationale, and the build

Phase 1 deliverable. Target: 1.0 B trained tokens (Β§2 band 900M–1.1B; frozen number confirmed at Gate 3). Every repo id, size, date and licence below was read live on 2026-09-19 from the Hub card, the /api/datasets/<id> metadata endpoint, or parquet footers over HTTP β€” not from memory. Where a card and the stored data disagree, that is recorded rather than smoothed over.

1. Selection rules applied, operationally

  • Β§3.4 English only. Every source is either English-config-scoped (20250702.en, eng_Latn, sample/10BT) or carries a per-row language field we filter on. Code and math are allowed and are deliberately included.
  • Β§3.5 No raw crawls. Excluded as a class: allenai/c4, HuggingFaceFW/fineweb, Skylion007/openwebtext, tiiuae/falcon-refinedweb. What is admitted instead is quality-scored (fineweb-edu carries a per-row classifier score/int_score), curated/rewritten (finewiki, cosmopedia), synthetic-textbook (finephrase), or openly-licensed domain corpora (common-pile/*, stack-v3-train). Common-Crawl appears only as the provenance of a classifier-filtered subset, never as the selection criterion.
  • Β§3.3 No contamination. Two mechanisms, not one assertion:
    1. Prefer sources with published decontamination. HuggingFaceTB/finemath documents 13-gram removal against GSM8K, MATH, MMLU and ARC test sets, with a public audit log (HuggingFaceTB/finemath_contamination_report).
    2. Reject sources whose own cards validate on the eval suite. That is why openbmb/UltraData-Math is dropped entirely despite being attractive on paper (apache-2.0, 170B+ tokens): its card lists MMLU/ARC-E/ARC-C/HellaSwag/PIQA/Winogrande/GSM8K/MATH500 as validation targets and its L3 tiers are synthetic exam-shaped text. Its en+zh mixing would also need a language filter. Similarly excluded: nvidia/Nemotron-Math-v2 (AoPS/SE solution traces, i.e. adjacent to GSM8K's own source pool, and solution-formatted in a way that quietly steers toward instruction data).
    3. The mechanical overlap audit of Β§5 runs against validation/dev material only; test splits stay untouched until Phase 6.

2. The mix: 1.31 B raw β†’ ~1.05 B after dedup β†’ ~1.0 B trained

# Source (config) raw tokens GB to pull licence rationale, from corpus properties
1 HuggingFaceFW/fineweb-edu sample/10BT, keep int_scoreβ‰₯4 300M 1.1 ODC-BY Classifier-scored general English; the only cheap source of broad register coverage. Score-gating is a quality filter, which Β§3.5 permits, and it is the source's own documented axis.
2 HuggingFaceFW/finepdfs-edu data/eng_Latn 150M 0.63 ODC-BY + CC ToU Globally-deduplicated long-form educational PDFs. Long document structure teaches coherence over >1k tokens, which chunked web prose does not. eng_Latn only (69 scripts exist).
3 HuggingFaceTB/cosmopedia stanford+openstax+khanacademy+auto_math_text 130M 0.33 Apache-2.0 Synthetic textbooks: causal and explanatory chains ("because", "therefore", worked steps). This is the register that builds science-QA and multi-step reasoning ability. Per-config GB measured, not estimated.
4 HuggingFaceTB/cosmopedia wikihow 30M 0.05 Apache-2.0 Procedural how-to steps β†’ ordered physical-world inference. Small slice, specific job.
5 HuggingFaceFW/finewiki data/enwiki 80M 0.43 CC-BY-SA-4.0 Link-resolved, rewritten encyclopedic prose; definitional and comparative sentence forms that factual/recall-style prompts reward. Heaviest GB/token in the mix, hence capped.
6 omarkamali/wikipedia-monthly 20250702.en 60M 0.03 CC-BY-SA-4.0 Cheapest encyclopedic tokens available (0.5 GB/Btok) and current; drop the raw_mediawiki column or it dominates the byte budget.
7 HuggingFaceTB/finemath finemath-4plus 130M 0.25 ODC-BY Best math density per GB (1.9) and the only candidate with a published eval-set decontamination report.
8 open-web-math/open-web-math 70M 0.15 ODC-BY (card body) Forum/webbook math in natural prose, incl. LaTeX. Complements #7's textbook skew with multi-turn human reasoning text. Oldest source here (2023-10) β€” cited as such.
9 HuggingFaceFW/finephrase tutorial + faq 90M 0.31 ODC-BY Rewritten into stepwise-tutorial and question/answer registers β€” forms largely absent from organic prose.
10 HuggingFaceFW/finephrase table 30M 0.10 ODC-BY Tabular→narrative realisation: reading quantities, units and row/column structure.
11 HuggingFaceCode/stack-v3-train, license_type=="permissive", per-language cap 130M 0.52 ODC-BY + per-file SPDX Code for compositional syntax and deterministic multi-step logic. Must filter to permissive licences (constraint 4) and cap XML/HTML/JSON/JS, which the card notes are byte-dominant.
12 SimpleStories/SimpleStories 40M 0.07 MIT Syntactically simple long-range narrative. Disproportionately useful below ~200M params, where the model otherwise never sees a consistent referent across a full context window.
13 common-pile/libretexts + common-pile/arxiv_abstracts 40M 0.08 CC0 / per-doc PD/CC-BY Openly-licensed OER textbooks and scientific abstracts; cleanest licence hygiene in the pool (arXiv metadata is CC0).
total raw 1,310M β‰ˆ4.35 GB

Headroom, and why it is sized like this. Β§Phase 2 requires slack above the target for dedup loss and a held-out split. Cross-source near-dedup is expected to cost 12–18 % (FinePhrase's four configs are rewrites of the same 338.7M source documents, so they must be deduplicated against the shared id/url or they buy topic echo rather than knowledge; #2/#5/#6 overlap on encyclopedic topics; #1/#2 overlap on educational web content). 1,310M Γ— 0.85 β‰ˆ 1,114M, minus a 2 % validation holdout (β‰ˆ22M) β†’ **1.09 B available**, which brackets the 1.0 B target and leaves room to land at 0.9 B without re-sourcing if Gate 3's throughput measurement forces the pre-registered fallback.

Sits comfortably inside the instance. β‰ˆ4.35 GB of shards + ~2 GB of Arrow/tokeniser working overhead + 2 GB of final token store β‰ͺ the measured 19.5 GB /kaggle/working. Every source here is pullable as individual 270 MB–3 GB shards, so the working directory is never the binding constraint β€” provided the build pulls shard-by-shard and deletes after tokenising, which Β§4 requires it to do.

3. Excluded, with reasons (so nobody re-litigates from scratch)

excluded why
allenai/c4, HuggingFaceFW/fineweb, Skylion007/openwebtext, tiiuae/falcon-refinedweb raw/unfiltered crawls β†’ Β§3.5
openai/gsm8k, cais/mmlu, allenai/ai2_arc, hellaswag, piqa, winogrande, truthfulqa the eval tasks themselves β†’ Β§3.3. Also: never glob a directory that could catch these.
nvidia/Nemotron-CC*, -Math-v1, -Code-v1 gated: manual; anonymous card and file reads return "Access to dataset … is restricted" β†’ token counts unverifiable and republication status unclear
openbmb/UltraData-Math card validates on the eval suite; exam-shaped L3 tiers; en+zh mixed
nvidia/Nemotron-Math-v2 solution-formatted AoPS/SE traces, GSM8K-adjacent; reads as instruction data
allenai/peS2o ~11 GB per Btok by measurement β€” worst ratio in the pool for 42 B tokens we do not need
HuggingFaceTB/smollm-corpus cosmopedia-v2 duplicate role with #3; also a documented card-vs-data mismatch (card 17.8 B vs ~31.7 B implied by stored token_length)
HuggingFaceFW/clean-wikipedia exists but README body is literally "Please see FineWiki instead" β†’ deprecated
HuggingFaceTB/smollm-corpus python-edu not a corpus. Verified schema is blob_id, repo_name, path, length_bytes, score, int_score β€” a filter list pointing at gated bigcode/the-stack-v2, with no text column. Do not budget code tokens from it.
bigcode/the-stack-v2-dedup, codeparrot/github-code not verified this session; the-stack-smol is gated: auto
TinyLlama 32k tokenizer Llama-2-derived; redistribution licence status unverified β†’ excluded from both tokenizer and any data
Cion-lab/Pretraining-10B-v1 (pre-existing in the account) composition/mix unknown, token counts estimated with the Qwen2.5 tokenizer, provenance not inspectable β†’ cannot be cleared under Β§3.5 or Β§3.3

4. Card-vs-data warning

Two sources disagree with themselves: cosmopedia-v2 (17.8 B card vs ~31.7 B stored) and finephrase (486 B "completion tokens" card vs ~1.4 B implied by stored rows Γ— inherited token_count). FinePhrase's token_count/score columns describe the source document, not the rewrite. Therefore: no token number in this project is ever taken from a card. Counts come from running our own tokenizer over the bytes, which is also what the Gate 2 manifest will record.

5. Build pipeline (Phase 2 design, on Kaggle CPU)

  1. Pull shard-by-shard, delete after consuming. 73–89 MB/s measured (35–89 across runs), so 4.35 GB arrives in minutes; the constraint is peak residency, not bandwidth. Never materialise the mix.
  2. Normalise + filter. Language filter on the row-level field where present; int_scoreβ‰₯4 for fineweb-edu; license_type=="permissive" + per-language cap for stack-v3; drop raw_mediawiki.
  3. Dedup, memory-bounded. Exact: 64-bit hash of the normalised document. Near: MinHash + LSH over 13-gram shingles (matching finemath's documented practice), processed in bands against a disk-backed index, because the box is a 30 GiB cgroup with no swap β€” a billion sketches in one dict is how a free-CPU job dies at hour six.
  4. Contamination audit (Β§3.3). Mechanical 13-gram overlap of the surviving mix against benchmark validation/dev material only β€” including MMLU's dev split specifically, since those are the few-shot exemplars the harness will use, so overlap there is contamination by construction. A script computes and writes counts per source; no item text is read, displayed or summarised by the agent. Any source failing a stated threshold is dropped or cleaned; result written up in docs/. Test splits are not opened until Phase 6.
  5. Tokenise with the frozen SmolLM2 BPE (49,152 β€” Β§3 of 01-plan). Measured cost on this shape: 660,943 tok/s across 4 cores β†’ 1.31 B tokens in ~0.55 h. Not a bottleneck; process-shard rather than thread-shard, since rayon scaling saturated at 2.7Γ— on 4 cores.
  6. Shard for exact positional resumption (Β§Phase 2 requirement). Concatenate token ids into fixed-size uint16 shards (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes): 40 M tokens = 80 MB per shard, ~28 shards, ~2 GB total. Then a resume position is literally (shard_index, token_offset) β€” one integer pair, memmap'd, so skip_first_batches on a replaying Trainer becomes an index advance rather than a decode. This is the property that makes Β§3.1's "same data, same position" checkable instead of hoped for.
  7. Hold out validation: ~22 M tokens sampled proportionally across sources, never whole documents (whole-document holdout leaks topic statistics and makes PPL optimistic), written as its own shard so validation PPL is bit-reproducible at Gate 4.
  8. Publish as public ounce100m-mix-v1 with a manifest recording: per-source and total token counts, shard count and size, tokenizer identity (repo + revision + sha256 of tokenizer.json), per-shard sha256, the dedup parameters, the audit counts, and the cursor semantics of step 6.

6. Licensing the derived mix

#5 and #6 are CC-BY-SA-4.0 (share-alike), and wikimedia additionally carries GFDL. A published derivative containing them must be offered under SA with attribution and a change log. That is compatible with Β§2's "repos are public", so the plan is: license the mix CC-BY-SA-4.0 and ship an ATTRIBUTION.md listing every source, its licence, and the transformations applied. ODC-BY sources (#1, #2, #7–#10) require notices kept and changes documented; stack-v3-train requires the per-file SPDX filter of Β§2 step 11 so we never republish a non-permissive file.

If a fully permissive mix is later required, drop #5 and #6 (140 M raw) and backfill from #1/#3. The trade is a small loss of encyclopedic density for a licence with no reciprocal obligation β€” recorded here so the choice is available and deliberate rather than discovered under time pressure.

7. Gate 2 checklist, restated as things that must be true

  • dataset live on the Hub as a public ounce100m-* repo
  • token count verified against the manifest by an independent recount, not by the builder's own number
  • contamination audit passed and written up (per-source overlap counts, validation/dev only)
  • validation split reserved, its token count recorded
  • shard layout proven small enough that a training job reads a few shards at a time and never materialises the mix β€” demonstrated by a job that reports peak /kaggle/working usage while consuming shards, not asserted from the arithmetic above
  • exact-position cursor (shard_index, token_offset) demonstrated to be reproducible

8. Risks

  1. Dedup is the only stage with real wall-clock cost and the only one that can exhaust RAM. Design for bands + disk; measure on one source before all thirteen.
  2. Unauthenticated Hub rate limits. Jobs currently read anonymously and the API warns about it; at ~50 shard downloads that is probably fine, but 429s mid-build are plausible. Mitigation is the mounted-credential path (Β§8 of 01-plan), which is being verified now.
  3. Card token counts are unreliable (Β§4) β€” every number in this file is either measured over HTTP or will be replaced by our own tokeniser's count at Gate 2.
  4. finepdfs-edu adds Common Crawl's Terms of Use on top of ODC-BY. Included because the subset is classifier-selected and globally deduplicated rather than a dump, but it is the least clean licence admission in the mix and the easiest to cut if the attribution review gets nervous.
  5. Two sources are rewrites/synthetic (#3, #4, #9, #10). That is permitted explicitly by Β§3.5, but a 34 % synthetic share can produce fluent-but-homogeneous prose; the proportions above keep organic English (sources 1, 2, 5, 6, 8, 12, 13) in the majority, and that balance is a deliberate choice worth revisiting only with evidence.

9. Phase 2 step 1 ran β€” Β§2 corrected by measurement

dodosoomro/ounce100m-p2-source-inventory (CPU, 116 s, 1,500 rows per source, tokenised with the frozen SmolLM2 tokenizer: vocab 49,152 confirmed, decode round-trips, 5.88 chars/token on a formal probe sentence). 13 of 18 entries resolved; 5 failed on config names, and two sources turned out to be unusable as described. Β§2's estimates were right in aggregate (predicted β‰ˆ4.35 GB, measured 4.46 GB for the 570 M tokens that resolved) but wrong in specifics, so the specifics are replaced rather than averaged.

9.1 Config-name corrections (the datasets-server truth, from the error messages themselves)

Β§2 said correct value note
fineweb-edu sample/10BT sample-10BT hyphen, not a path. Full set: default, sample-10BT, sample-100BT, sample-350BT, plus per-dump CC-MAIN-2025-05/08/13/18/21/26…
finemath finemath4plus finemath-4plus also finemath-3plus, infiwebmath-3plus, infiwebmath-4plus
finewiki data/enwiki config en configs are plain language codes (ab, ace, af, …), not directory paths
wikipedia-monthly 20250702.en 20260101.en the 2025-07 config is gone; newer dumps now exist (20260101.*), which is better for currency anyway
stack-v3-train python default only there are no per-language configs. Language is a row field, and rows are whole repositories with a nested files struct

Every one of these would have been a mid-build failure. Recorded in docs/00-platform-notes.md style: the card's own naming is not an API β€” load_dataset's error is.

9.2 Measured yield per document β€” this is what the token arithmetic actually runs on

source tokens/doc chars/doc chars/token non-English-flagged rows verdict
finepdfs-edu eng_Latn 1,321 14,352 3.98 128/1500 = 8.5 % keep, must add our own language filter
cosmopedia stanford 932 4,917 5.30 15/1500 = 1.0 % keep
cosmopedia wikihow 907 4,541 4.97 0.7 % keep
cosmopedia openstax 765 3,820 4.93 0.8 % keep
cosmopedia auto_math_text 689 2,811 4.14 0.8 % keep
cosmopedia khanacademy 677 3,047 4.26 1.1 % keep
open-web-math default 1,182 6,795 3.31 112/1500 = 7.5 % keep, needs en filter; low chars/token = LaTeX density, which is the point
finephrase tutorial 779 4,903 4.58 1.1 % keep. Column is text, not completion
finephrase faq 744 4,965 4.60 1.1 % keep
finephrase table 715 4,772 4.59 1.1 % keep
SimpleStories 285 1,260 4.33 0.5 % keep, but it is 4Γ— the row count for the same tokens
common-pile/arxiv_abstracts 185 791 4.31 2.2 % keep. Column is text, not raw_content
common-pile/libretexts 75 964 2.42 1396/1500 = 93 % DROP

libretexts is disqualified, and the reason matters. Its text field is not prose: 93 % of sampled rows failed a naive English test, chars/token collapsed to 2.42, and the first example is a bare https://math.libretexts.org/Bookshelves/... URL. The Common-Pile schema puts real content in metadata/raw_content behind a per-source convention, so the column we assumed is an identifier field. Rather than reverse-engineer four sub-corpora's schemas for 10–20M tokens of marginal value, it is cut. arxiv_abstracts from the same publisher did work (its text is real abstract prose), so this is a per-dataset fact and not a Common-Pile-wide one.

9.3 Two open problems this created, and how they are being handled

  1. Code no longer has a cheap path. With no per-language config, getting 60–130M tokens of Python/JS from stack-v3-train means streaming whole-repo rows, opening a nested files struct, and filtering on a row field β€” for a source that concentrates in XML/HTML/JSON anyway. Code is not mandatory (Β§3.4 makes proportions our call), and at 100M params / 1B tokens the marginal value is genuinely unclear. Action: a dedicated mini-probe before committing β€” measure rows-sampled-per-Python-repo and the nested schema β€” and hold code at ≀60M in the plan until that returns. If it is awkward, drop it; that is a defensible Β§3.4 choice, not a failure.
  2. eng_Latn and open-web-math are not actually English-only (8.5 % and 7.5 % of sampled rows fail a crude English test; the finepdfs sample included Cyrillic inline in a single "English" document). Β§3.4 is English-only, so a language filter is now a required pipeline stage, not a courtesy β€” applied per row on our side, on top of whatever the upstream config claims. finepdfs carries full_doc_lid / per_page_languages columns to do this cheaply; the general fallback is language_score where present and a fastText-style gate where not.

9.4 Revised totals

Confirmed-and-usable now: 570M target tokens across 13 resolved configs, β‰ˆ4.5 GB, at measured yields of roughly 909k documents β€” plus fineweb-edu/finewiki/wikipedia-monthly/finemath recoverable with the corrected config names (Β§9.1), which restores ~570M of the planned 1,310M. Net effect on the plan: the shape survives, the line items do not. Β§2's per-source token targets stay as targets; Β§2's config strings and two column names are superseded by this section.

Directly measured consequence for the build: 909k documents for 570M tokens β‰ˆ 627 tokens/document average, so the mix is ~1.6–2.1M documents for 1.0–1.3B tokens. That is the number the dedup design has to be sized against β€” under 2.1M documents, a MinHash signature table comfortably fits the 30 GiB RAM, which materially de-risks the one stage Β§8 called the real cost.

10. Builder implemented, and one blocking flaw found and fixed

code/build/build_mix.py (mirrored in Cion-lab/ounce100m-code) is the Phase 2 builder. It stages each source separately (stream β†’ normalise β†’ gates β†’ dedup β†’ tokenise) and then merges the staged shards into the final mixed shards by largest-deficit round-robin over documents.

E-016, and why the smoke test was mandatory. The first version consumed one source to its target and only flushed a shard when the buffer filled. Since fineweb-edu's target alone is ~37 shards of 8M tokens, the first third of the mix would have been nothing but web text, then textbooks, then math β€” a curriculum drift across the run that no LR schedule can repair, and one the builder itself reported as success (checksums, headers, token counts and cursor probes were all clean). The verifier caught it as distinct_sources_per_shard_min: 1. The shuffle that was supposed to fix ordering only permuted documents inside a single-source buffer. After the two-stage rewrite, the same smoke run reports min = max = median = 4 sources in every shard, 12,001,386 tokens recounted from the bytes, 0 out-of-range ids, 40/40 cursor probes and 40/40 document-boundary reconstructions correct.

Shard format (shared by staged and final, so document boundaries survive to the trainer):

header  : <IIII  = vocab, n_docs, n_tokens, flags        16 B
offsets : uint32 Γ— (n_docs + 1), cumulative, [0] = 0
ids     : uint16 Γ— n_tokens        (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes)
doc i   = ids[offsets[i] : offsets[i+1]]

Target 8M tokens/shard (16 MB, ~150 shards, ~2.5 GB for the mix). The resume cursor is (shard_index, token_offset) β€” one integer pair, which is what makes Β§3.1's "same data, same position" checkable and what keeps Trainer's replaying skip_first_batches cheap on a map-style read.

Interruption safety, which was not optional. /kaggle/working is destroyed when a session ends, and a full staging pass is hours of streaming. So each source is published to Cion-lab/ounce100m-mix-stage (the stage repo, distinct from the final ounce100m-mix-v1) the moment it completes, with a record.json holding its shard list, token count and drop statistics. A new session rebuilds its stage bookkeeping from the Hub and skips every source already published; only the source in flight at the moment of the kill is re-staged. code/build/smoke_hub_resume.py rehearses this by wiping the local tree between stage and merge, because "merge-only with an empty state and no restore path" is a failure mode that exits with code 0 while emitting an empty mix.

Revised raw total: 1.23 B tokens = 23 % headroom over the 1.0 B trained target (raised fineweb-edu to 340M, finepdfs-edu to 180M, finemath to 160M, because Β§9's target edits had quietly pulled the plan from 1.31 B down to ~1.13 B β€” about 11 % headroom, which is not headroom once a 12–18 % dedup loss is applied). 16 sources, no code source yet (Β§9.3).

Gates the build must satisfy before Gate 2 is claimed, as assertions in verify_mix.py rather than prose: every shard sha256 matches the manifest; no trailing bytes; offsets well-formed; zero out-of-range ids; total tokens recounted from bytes equals the manifest; and β‰₯ half the distinct sources present in every shard. The last one exists because the first build passed everything else.

11. Interruption safety proven, and the build that this authorises

A CPU session's /kaggle/working is destroyed when it ends. Everything below was measured, not assumed (dodosoomro/ounce100m-p2-cold-resume-rehearsal v6, 17:35Z, 55 s wall-clock, pinned REV 428c5234).

Step Observed
1. stage two sources, publish each 2,400,867 tokens staged; record.json + shard-0000.bin per source on the Hub
2. wipe the local tree (root_exists: false, stage_dir_exists: false) faithful simulation of a session kill, not a soft restart
3. --merge-only with the Hub as the only state hub: restored 2 staged source(s) (2,400,867 tokens) β†’ merged shard 2,352,747 tok / 5,815 docs, 2/2 sources in it, val shard 48,120 tok
4. verify_mix.py over the restored mix PASS: true, sha_ok 1/1, out_of_range_ids: 0, trailing_bytes: 0, cursor probes 40/40 boundaries reconstructed, 0 mismatches, recount_matches_manifest: true
5. verdict + self-cleanup COLD_RESUME_PASSED=True, cleanup=deleted, throwaway repo confirmed gone from the namespace

v4 of the same rehearsal had reported false over an identically-working pipeline: a stray print("shards=", …) inside the BUILD_JSON block made the parent's json.loads fail and a helper fold that into None (memory/ERRORS.md E-019, and the report-shape recurrence E-023). The fix was to the harness, plus a (obj, why_null) return and a manifest.json fallback so a verdict can never again be silently null.

The build now running (dodosoomro/ounce100m-p2-full-mix-build, 17:38Z, CPU, sessionTimeoutSeconds 43200, pinned REV 98e49560): build_mix.py --root /kaggle/working/mixroot --hub-repo Cion-lab/ounce100m-mix-stage with defaults β€” stage cap 1.31 B tokens, per-source targets summing to 1.23 B, 8 M tokens/mixed shard, VAL_STRIDE 50 holdout, shuffle seed 20260919. Each source publishes as it completes, so an interruption costs at most one source's streaming time, and a re-launch restores the finished ones from the Hub (step 3 above is that exact path).

Audit semantics, hardened the same way. audit_contamination.py compares 13-token windows against {train, validation, dev} of the eight tasks β€” never test β€” and emits counts only. It now also reports shards_missing, and its verdict is composite: AUDIT_PASSED requires every manifest shard present and read, > 0 documents and grams measured, β‰₯ 6 readable reference sets, zero reference errors, and zero overlap. Exit 6 means overlap found (a real finding); exit 7 means the audit did not measure (a bug). A zero-overlap report over shards that were not on disk is no longer expressible.

Publishing ships the tokenizer. publish_mix.py copies tokenizer.json into the dataset root, re-hashes it and aborts with exit 8 unless it matches the sha256 recorded in the manifest β€” so the published ids and the published tokenizer are provably the pair that produced them (Β§3.12).

12. Gate 2 β€” the record, including where it failed and what that failure was worth

Merge and verify: passed, twice, reproducibly. From the 15 staged sources (Cion-lab/ounce100m-mix-stage, 1,150,015,841 tokens staged):

field value
merged shards 141 (each drawn from all 15 sources)
train tokens 1,127,163,952
held-out validation 22,851,889 in 12 shards (every 50th merged document)
documents 1,078,807
per-source share 75,144,263 tokens each, equal by construction
merge time 16.4 s (restore from the Hub included: 123.8 s)
verify_mix PASS: true β€” sha256 141/141, headers 141/141, 0 bad offsets, 0 out-of-range ids, 0 trailing bytes, recount matches manifest, 40/40 cursor probes, distinct_sources_per_shard_min = 15 (floor was 7)
finewiki__en never staged. Gate 2's own diagnostic showed the columns present and first_row_ok: true with a truncated ImportError in the loading script. 15 sources at 1.127 B is above the 1.0 B target with 12.7 % headroom, so this is a composition note, not a blocker

Then the audit refused to publish, and it was right to. audit_contamination exited 7:

task 13-gram hits reference size (verified against the dataset server)
mmlu 904,288 dev 285 + validation 1,531 items β‰ˆ 149k grams
gsm8k 185,600 train 7,473 items = 360,656 grams
hellaswag 1,852 train 39,905 + validation 10,042
arc_easy 358 train 2,251 + validation 570
arc_challenge 214 train 1,119 + validation 299
piqa not measured the loader yielded 0 of 16,113 train rows without raising
winogrande (not in the visible excerpt) train 40,398 + validation 1,267
truthfulqa (repo name was wrong) validation 817, multiple_choice

expected_false_positives for the whole scan was β‰ˆ0.0001, so these are matches, not collisions. The first reading I wrote in STATE.md β€” "GSM8K's whole train split is in the mix" β€” was an over-claim: hits are counted per mix gram, so a small number of frequently repeated reference grams can produce a large count, and 185,600 hits against 360,656 distinct GSM8K grams means "up to half of its gram vocabulary occurs somewhere in 1.1 B tokens". The columns that settle severity are mix_documents_majority_hit (documents whose grams are more than half matched β†’ verbatim inclusion) and source_hit_document_fraction (which source, and how much of it) β€” and both are only computable on the fixed audit, which is what Gate 2c/2d exist to produce.

Why upstream did not catch it, verified rather than assumed: FineMath's card states it removed 13-gram overlap against the test sets of GSM8k, MATH, MMLU and ARC; Cosmopedia's describes a 10-gram + difflib pass against test benchmarks; OpenWebMath documents only SimHash self-deduplication with no benchmark claim. validation/dev/train overlap was therefore never in scope for any of them β€” and validation is the split lm-eval scores HellaSwag, PIQA, Winogrande and TruthfulQA on (docs/05-eval-plan.md Β§7).

Gate 2 remained unsigned at the end of this stage. Cion-lab/ounce100m-mix-v1 did not exist yet. The path to it: attribution β†’ drop the offending documents by mask (or a source if attribution says it is broadly soaked) β†’ re-merge with --exclude-dir β†’ re-audit until AUDIT_PASSED with all 8 tasks covered β†’ publish_mix, which now refuses to ship unless the audit passed, covered all 8, scanned the val shards, and its mix_bytes_sha256 matches the exact manifest being published. E-030 explains why the first filtered re-merge could not work and what was changed.

12.4 Resolution β€” the filtered mix, measured (Gates 2b–2f, 2026-09-19)

Attribution first, always. Gate 2b/2c scanned the staged shards per source, which is the only way to say where contamination lives rather than how much the merged mix contains. Of 1,955,832 staged documents, 10,425 (0.533 %) carried at least one matched 13-gram and 0 were majority-matched β€” shared spans, not wholesale inclusion. The count concentrated in five sources: finemath-4plus 2,688 documents, fineweb-edu 1,043, cosmopedia auto_math_text 568, finepdfs-edu 420, open-web-math 339, with the remaining ~5,367 spread over ten more. Per-source hit counts summed to ~50,700 while the merged scan reported 2,096,540 for essentially the same gram volume; the discrepancy is unreconciled and recorded as E-032c rather than smoothed over. It changes no decision: the exclusion is per-document and the post-filter overlap is zero either way.

The exclusion, and what it cost. audit_contamination --write-filter emitted one .u8 mask per source (1 byte per staged document, ~700 KB total) marking documents that contain a matched window. Re-merging with --exclude-dir dropped 5,878 documents = 17,366,967 tokens (1.5 % of the mix) and the re-audit of the result gave overlap_total = 0 with 8/8 benchmark tasks covered, the held-out validation shards scanned clean, and MEASURED: true. Final mix: 1,109,714,831 training tokens in 141 shards plus 22,851,889 held-out validation tokens β€” still 11.0 % above the 1.0 B target, so the pre-registered 0.9 B fallback (D-012) did not fire and the contamination fix cost nothing that the token budget had to pay for.

Reference coverage is now a precondition, not an assumption. Gate 2b found that piqa's train and validation splits returned 0 rows from the datasets library inside the Kaggle image while the server reports 16,113 + 1,838 β€” the first audit had scored that as "clean". The audit now fetches each split's server-side row count (/size), refuses any split it read to under 90 % of it, retries through the HTTP /rows endpoint, and marks the whole report MEASURED: false if a loader still falls short. TruthfulQA's answer text also sits under keys (mc1_targets, mc2_targets, abstractions) the first version never read. All reference splits used are train/validation/dev; no test split was opened (Β§3.3).

Why publication took three attempts (E-033). Gates 2e and 2f both built and audited the mix correctly and then died in the uploader: publish_mix sent one commit per file, and the Hub began rejecting every request β€” including 30-byte files β€” after ~139 commits. That is a commit-rate ceiling, not a network fault, and the "resumable" re-send was itself broken because RepoFile.lfs is a dict and getattr(ent.lfs, "sha256", None) is therefore None for every shard, i.e. nothing looked published. The uploader now reads lfs["sha256"] or lfs["oid"], sends the remaining files as one upload_folder commit, and after the commit re-lists the repo and re-hashes, exiting 11 if a path is absent and 12 if a remote digest disagrees with the local bytes. Cion-lab/ounce100m-mix-v1 held 142 of 173 files when Gate 2g re-ran the same pipeline with it.

Gate 2 is signed (dodosoomro/ounce100m-gate-2g-publish-mix-v1, complete 2026-09-20T00:31Z, CPU, zero GPU quota). The rebuilt mix reproduced Gate 2e's token total exactly β€” 1,109,714,831 β€” which is the determinism claim in Β§11 tested a second time on a different session with the same inputs. audit.json on the Hub reports overlap_total = 0, mix_documents_with_any_hit = 0, majority_hit_docs = 0, expected_false_positives = 0.0, tasks_covered = all eight, val_hits = 0 across 21,898 held-out documents, and AUDIT_PASSED, MEASURED, COVERED_ALL_SHARDS all true, bound to these exact bytes by mix_bytes_sha256 = 732eec7f19e77f36…. verify_mix exited 0. The batched uploader landed all 174 files (139 train shards, 12 val shards, 16 filter/ objects, manifest.json, audit.json, tokenizer.json, README.md, ATTRIBUTION.md, build_state.json, .gitattributes), and four train plus two validation shards were then read back from the anonymous resolve endpoint and matched the manifest's sha256 β€” so the mix is public and serving, not merely listed (Β§3.13).

One editorial note on this section: Β§12.2's first draft treated overlap_total as a count of contaminated documents. It is a count of matched grams, which is why a 41Γ— difference between the per-source and merged hit totals could coexist with a 1.3 % agreement on document counts (E-032c). The action taken β€” drop the documents, re-audit to zero β€” was unaffected, but the number in the record was read wrongly for two gates, and the fix was to write the density columns (mix_documents_with_any_hit, mix_documents_majority_hit) and quote those when judging severity.