Download docs/02-mix-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 36 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/02-mix-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/02-mix-plan.md
-
curl -L -o 02-mix-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/02-mix-plan.md
02 β Mix plan: sources, counts, rationale, and the build
Phase 1 deliverable. Target: 1.0 B trained tokens (Β§2 band 900Mβ1.1B; frozen number confirmed at
Gate 3). Every repo id, size, date and licence below was read live on 2026-09-19 from the Hub card, the
/api/datasets/<id> metadata endpoint, or parquet footers over HTTP β not from memory. Where a card and
the stored data disagree, that is recorded rather than smoothed over.
1. Selection rules applied, operationally
- Β§3.4 English only. Every source is either English-config-scoped (
20250702.en,eng_Latn,sample/10BT) or carries a per-row language field we filter on. Code and math are allowed and are deliberately included. - Β§3.5 No raw crawls. Excluded as a class:
allenai/c4,HuggingFaceFW/fineweb,Skylion007/openwebtext,tiiuae/falcon-refinedweb. What is admitted instead is quality-scored (fineweb-educarries a per-row classifierscore/int_score), curated/rewritten (finewiki,cosmopedia), synthetic-textbook (finephrase), or openly-licensed domain corpora (common-pile/*,stack-v3-train). Common-Crawl appears only as the provenance of a classifier-filtered subset, never as the selection criterion. - Β§3.3 No contamination. Two mechanisms, not one assertion:
- Prefer sources with published decontamination.
HuggingFaceTB/finemathdocuments 13-gram removal against GSM8K, MATH, MMLU and ARC test sets, with a public audit log (HuggingFaceTB/finemath_contamination_report). - Reject sources whose own cards validate on the eval suite. That is why
openbmb/UltraData-Mathis dropped entirely despite being attractive on paper (apache-2.0, 170B+ tokens): its card lists MMLU/ARC-E/ARC-C/HellaSwag/PIQA/Winogrande/GSM8K/MATH500 as validation targets and its L3 tiers are synthetic exam-shaped text. Its en+zh mixing would also need a language filter. Similarly excluded:nvidia/Nemotron-Math-v2(AoPS/SE solution traces, i.e. adjacent to GSM8K's own source pool, and solution-formatted in a way that quietly steers toward instruction data). - The mechanical overlap audit of Β§5 runs against validation/dev material only; test splits stay untouched until Phase 6.
- Prefer sources with published decontamination.
2. The mix: 1.31 B raw β ~1.05 B after dedup β ~1.0 B trained
| # | Source (config) | raw tokens | GB to pull | licence | rationale, from corpus properties |
|---|---|---|---|---|---|
| 1 | HuggingFaceFW/fineweb-edu sample/10BT, keep int_scoreβ₯4 |
300M | 1.1 | ODC-BY | Classifier-scored general English; the only cheap source of broad register coverage. Score-gating is a quality filter, which Β§3.5 permits, and it is the source's own documented axis. |
| 2 | HuggingFaceFW/finepdfs-edu data/eng_Latn |
150M | 0.63 | ODC-BY + CC ToU | Globally-deduplicated long-form educational PDFs. Long document structure teaches coherence over >1k tokens, which chunked web prose does not. eng_Latn only (69 scripts exist). |
| 3 | HuggingFaceTB/cosmopedia stanford+openstax+khanacademy+auto_math_text |
130M | 0.33 | Apache-2.0 | Synthetic textbooks: causal and explanatory chains ("because", "therefore", worked steps). This is the register that builds science-QA and multi-step reasoning ability. Per-config GB measured, not estimated. |
| 4 | HuggingFaceTB/cosmopedia wikihow |
30M | 0.05 | Apache-2.0 | Procedural how-to steps β ordered physical-world inference. Small slice, specific job. |
| 5 | HuggingFaceFW/finewiki data/enwiki |
80M | 0.43 | CC-BY-SA-4.0 | Link-resolved, rewritten encyclopedic prose; definitional and comparative sentence forms that factual/recall-style prompts reward. Heaviest GB/token in the mix, hence capped. |
| 6 | omarkamali/wikipedia-monthly 20250702.en |
60M | 0.03 | CC-BY-SA-4.0 | Cheapest encyclopedic tokens available (0.5 GB/Btok) and current; drop the raw_mediawiki column or it dominates the byte budget. |
| 7 | HuggingFaceTB/finemath finemath-4plus |
130M | 0.25 | ODC-BY | Best math density per GB (1.9) and the only candidate with a published eval-set decontamination report. |
| 8 | open-web-math/open-web-math |
70M | 0.15 | ODC-BY (card body) | Forum/webbook math in natural prose, incl. LaTeX. Complements #7's textbook skew with multi-turn human reasoning text. Oldest source here (2023-10) β cited as such. |
| 9 | HuggingFaceFW/finephrase tutorial + faq |
90M | 0.31 | ODC-BY | Rewritten into stepwise-tutorial and question/answer registers β forms largely absent from organic prose. |
| 10 | HuggingFaceFW/finephrase table |
30M | 0.10 | ODC-BY | Tabularβnarrative realisation: reading quantities, units and row/column structure. |
| 11 | HuggingFaceCode/stack-v3-train, license_type=="permissive", per-language cap |
130M | 0.52 | ODC-BY + per-file SPDX | Code for compositional syntax and deterministic multi-step logic. Must filter to permissive licences (constraint 4) and cap XML/HTML/JSON/JS, which the card notes are byte-dominant. |
| 12 | SimpleStories/SimpleStories |
40M | 0.07 | MIT | Syntactically simple long-range narrative. Disproportionately useful below ~200M params, where the model otherwise never sees a consistent referent across a full context window. |
| 13 | common-pile/libretexts + common-pile/arxiv_abstracts |
40M | 0.08 | CC0 / per-doc PD/CC-BY | Openly-licensed OER textbooks and scientific abstracts; cleanest licence hygiene in the pool (arXiv metadata is CC0). |
| total raw | 1,310M | β4.35 GB |
Headroom, and why it is sized like this. Β§Phase 2 requires slack above the target for dedup loss
and a held-out split. Cross-source near-dedup is expected to cost 12β18 % (FinePhrase's four configs
are rewrites of the same 338.7M source documents, so they must be deduplicated against the shared
1.09 B available**, which brackets the 1.0 B target and leaves room to land at 0.9 B
without re-sourcing if Gate 3's throughput measurement forces the pre-registered fallback.id/url or they buy topic echo rather than knowledge; #2/#5/#6 overlap on encyclopedic topics;
#1/#2 overlap on educational web content). 1,310M Γ 0.85 β 1,114M, minus a 2 % validation holdout
(β22M) β **
Sits comfortably inside the instance. β4.35 GB of shards + ~2 GB of Arrow/tokeniser working
overhead + 2 GB of final token store βͺ the measured 19.5 GB /kaggle/working. Every source here is
pullable as individual 270 MBβ3 GB shards, so the working directory is never the binding constraint β
provided the build pulls shard-by-shard and deletes after tokenising, which Β§4 requires it to do.
3. Excluded, with reasons (so nobody re-litigates from scratch)
| excluded | why |
|---|---|
allenai/c4, HuggingFaceFW/fineweb, Skylion007/openwebtext, tiiuae/falcon-refinedweb |
raw/unfiltered crawls β Β§3.5 |
openai/gsm8k, cais/mmlu, allenai/ai2_arc, hellaswag, piqa, winogrande, truthfulqa |
the eval tasks themselves β Β§3.3. Also: never glob a directory that could catch these. |
nvidia/Nemotron-CC*, -Math-v1, -Code-v1 |
gated: manual; anonymous card and file reads return "Access to dataset β¦ is restricted" β token counts unverifiable and republication status unclear |
openbmb/UltraData-Math |
card validates on the eval suite; exam-shaped L3 tiers; en+zh mixed |
nvidia/Nemotron-Math-v2 |
solution-formatted AoPS/SE traces, GSM8K-adjacent; reads as instruction data |
allenai/peS2o |
~11 GB per Btok by measurement β worst ratio in the pool for 42 B tokens we do not need |
HuggingFaceTB/smollm-corpus cosmopedia-v2 |
duplicate role with #3; also a documented card-vs-data mismatch (card 17.8 B vs ~31.7 B implied by stored token_length) |
HuggingFaceFW/clean-wikipedia |
exists but README body is literally "Please see FineWiki instead" β deprecated |
HuggingFaceTB/smollm-corpus python-edu |
not a corpus. Verified schema is blob_id, repo_name, path, length_bytes, score, int_score β a filter list pointing at gated bigcode/the-stack-v2, with no text column. Do not budget code tokens from it. |
bigcode/the-stack-v2-dedup, codeparrot/github-code |
not verified this session; the-stack-smol is gated: auto |
TinyLlama 32k tokenizer |
Llama-2-derived; redistribution licence status unverified β excluded from both tokenizer and any data |
Cion-lab/Pretraining-10B-v1 (pre-existing in the account) |
composition/mix unknown, token counts estimated with the Qwen2.5 tokenizer, provenance not inspectable β cannot be cleared under Β§3.5 or Β§3.3 |
4. Card-vs-data warning
Two sources disagree with themselves: cosmopedia-v2 (17.8 B card vs ~31.7 B stored) and finephrase
(486 B "completion tokens" card vs ~1.4 B implied by stored rows Γ inherited token_count). FinePhrase's
token_count/score columns describe the source document, not the rewrite. Therefore: no token
number in this project is ever taken from a card. Counts come from running our own tokenizer over the
bytes, which is also what the Gate 2 manifest will record.
5. Build pipeline (Phase 2 design, on Kaggle CPU)
- Pull shard-by-shard, delete after consuming. 73β89 MB/s measured (35β89 across runs), so 4.35 GB arrives in minutes; the constraint is peak residency, not bandwidth. Never materialise the mix.
- Normalise + filter. Language filter on the row-level field where present;
int_scoreβ₯4for fineweb-edu;license_type=="permissive"+ per-language cap for stack-v3; dropraw_mediawiki. - Dedup, memory-bounded. Exact: 64-bit hash of the normalised document. Near: MinHash + LSH over 13-gram shingles (matching finemath's documented practice), processed in bands against a disk-backed index, because the box is a 30 GiB cgroup with no swap β a billion sketches in one dict is how a free-CPU job dies at hour six.
- Contamination audit (Β§3.3). Mechanical 13-gram overlap of the surviving mix against benchmark
validation/dev material only β including MMLU's
devsplit specifically, since those are the few-shot exemplars the harness will use, so overlap there is contamination by construction. A script computes and writes counts per source; no item text is read, displayed or summarised by the agent. Any source failing a stated threshold is dropped or cleaned; result written up indocs/. Test splits are not opened until Phase 6. - Tokenise with the frozen SmolLM2 BPE (49,152 β Β§3 of
01-plan). Measured cost on this shape: 660,943 tok/s across 4 cores β 1.31 B tokens in ~0.55 h. Not a bottleneck; process-shard rather than thread-shard, since rayon scaling saturated at 2.7Γ on 4 cores. - Shard for exact positional resumption (Β§Phase 2 requirement). Concatenate token ids into fixed-size
uint16shards (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes): 40 M tokens = 80 MB per shard, ~28 shards, ~2 GB total. Then a resume position is literally(shard_index, token_offset)β one integer pair, memmap'd, soskip_first_batcheson a replayingTrainerbecomes an index advance rather than a decode. This is the property that makes Β§3.1's "same data, same position" checkable instead of hoped for. - Hold out validation: ~22 M tokens sampled proportionally across sources, never whole documents (whole-document holdout leaks topic statistics and makes PPL optimistic), written as its own shard so validation PPL is bit-reproducible at Gate 4.
- Publish as public
ounce100m-mix-v1with a manifest recording: per-source and total token counts, shard count and size, tokenizer identity (repo + revision + sha256 oftokenizer.json), per-shard sha256, the dedup parameters, the audit counts, and the cursor semantics of step 6.
6. Licensing the derived mix
#5 and #6 are CC-BY-SA-4.0 (share-alike), and wikimedia additionally carries GFDL. A published
derivative containing them must be offered under SA with attribution and a change log. That is compatible
with Β§2's "repos are public", so the plan is: license the mix CC-BY-SA-4.0 and ship an
ATTRIBUTION.md listing every source, its licence, and the transformations applied. ODC-BY sources
(#1, #2, #7β#10) require notices kept and changes documented; stack-v3-train requires the per-file
SPDX filter of Β§2 step 11 so we never republish a non-permissive file.
If a fully permissive mix is later required, drop #5 and #6 (140 M raw) and backfill from #1/#3. The trade is a small loss of encyclopedic density for a licence with no reciprocal obligation β recorded here so the choice is available and deliberate rather than discovered under time pressure.
7. Gate 2 checklist, restated as things that must be true
- dataset live on the Hub as a public
ounce100m-*repo - token count verified against the manifest by an independent recount, not by the builder's own number
- contamination audit passed and written up (per-source overlap counts, validation/dev only)
- validation split reserved, its token count recorded
- shard layout proven small enough that a training job reads a few shards at a time and never
materialises the mix β demonstrated by a job that reports peak
/kaggle/workingusage while consuming shards, not asserted from the arithmetic above - exact-position cursor
(shard_index, token_offset)demonstrated to be reproducible
8. Risks
- Dedup is the only stage with real wall-clock cost and the only one that can exhaust RAM. Design for bands + disk; measure on one source before all thirteen.
- Unauthenticated Hub rate limits. Jobs currently read anonymously and the API warns about it; at
~50 shard downloads that is probably fine, but 429s mid-build are plausible. Mitigation is the
mounted-credential path (Β§8 of
01-plan), which is being verified now. - Card token counts are unreliable (Β§4) β every number in this file is either measured over HTTP or will be replaced by our own tokeniser's count at Gate 2.
finepdfs-eduadds Common Crawl's Terms of Use on top of ODC-BY. Included because the subset is classifier-selected and globally deduplicated rather than a dump, but it is the least clean licence admission in the mix and the easiest to cut if the attribution review gets nervous.- Two sources are rewrites/synthetic (#3, #4, #9, #10). That is permitted explicitly by Β§3.5, but a 34 % synthetic share can produce fluent-but-homogeneous prose; the proportions above keep organic English (sources 1, 2, 5, 6, 8, 12, 13) in the majority, and that balance is a deliberate choice worth revisiting only with evidence.
9. Phase 2 step 1 ran β Β§2 corrected by measurement
dodosoomro/ounce100m-p2-source-inventory (CPU, 116 s, 1,500 rows per source, tokenised with the
frozen SmolLM2 tokenizer: vocab 49,152 confirmed, decode round-trips, 5.88 chars/token on a formal
probe sentence). 13 of 18 entries resolved; 5 failed on config names, and two sources turned out to be
unusable as described. Β§2's estimates were right in aggregate (predicted β4.35 GB, measured 4.46 GB for
the 570 M tokens that resolved) but wrong in specifics, so the specifics are replaced rather than
averaged.
9.1 Config-name corrections (the datasets-server truth, from the error messages themselves)
| Β§2 said | correct value | note |
|---|---|---|
fineweb-edu sample/10BT |
sample-10BT |
hyphen, not a path. Full set: default, sample-10BT, sample-100BT, sample-350BT, plus per-dump CC-MAIN-2025-05/08/13/18/21/26β¦ |
finemath finemath4plus |
finemath-4plus |
also finemath-3plus, infiwebmath-3plus, infiwebmath-4plus |
finewiki data/enwiki |
config en |
configs are plain language codes (ab, ace, af, β¦), not directory paths |
wikipedia-monthly 20250702.en |
20260101.en |
the 2025-07 config is gone; newer dumps now exist (20260101.*), which is better for currency anyway |
stack-v3-train python |
default only |
there are no per-language configs. Language is a row field, and rows are whole repositories with a nested files struct |
Every one of these would have been a mid-build failure. Recorded in docs/00-platform-notes.md style:
the card's own naming is not an API β load_dataset's error is.
9.2 Measured yield per document β this is what the token arithmetic actually runs on
| source | tokens/doc | chars/doc | chars/token | non-English-flagged rows | verdict |
|---|---|---|---|---|---|
finepdfs-edu eng_Latn |
1,321 | 14,352 | 3.98 | 128/1500 = 8.5 % | keep, must add our own language filter |
cosmopedia stanford |
932 | 4,917 | 5.30 | 15/1500 = 1.0 % | keep |
cosmopedia wikihow |
907 | 4,541 | 4.97 | 0.7 % | keep |
cosmopedia openstax |
765 | 3,820 | 4.93 | 0.8 % | keep |
cosmopedia auto_math_text |
689 | 2,811 | 4.14 | 0.8 % | keep |
cosmopedia khanacademy |
677 | 3,047 | 4.26 | 1.1 % | keep |
open-web-math default |
1,182 | 6,795 | 3.31 | 112/1500 = 7.5 % | keep, needs en filter; low chars/token = LaTeX density, which is the point |
finephrase tutorial |
779 | 4,903 | 4.58 | 1.1 % | keep. Column is text, not completion |
finephrase faq |
744 | 4,965 | 4.60 | 1.1 % | keep |
finephrase table |
715 | 4,772 | 4.59 | 1.1 % | keep |
SimpleStories |
285 | 1,260 | 4.33 | 0.5 % | keep, but it is 4Γ the row count for the same tokens |
common-pile/arxiv_abstracts |
185 | 791 | 4.31 | 2.2 % | keep. Column is text, not raw_content |
common-pile/libretexts |
75 | 964 | 2.42 | 1396/1500 = 93 % | DROP |
libretexts is disqualified, and the reason matters. Its text field is not prose: 93 % of sampled
rows failed a naive English test, chars/token collapsed to 2.42, and the first example is a bare
https://math.libretexts.org/Bookshelves/... URL. The Common-Pile schema puts real content in
metadata/raw_content behind a per-source convention, so the column we assumed is an identifier field.
Rather than reverse-engineer four sub-corpora's schemas for 10β20M tokens of marginal value, it is cut.
arxiv_abstracts from the same publisher did work (its text is real abstract prose), so this is a
per-dataset fact and not a Common-Pile-wide one.
9.3 Two open problems this created, and how they are being handled
- Code no longer has a cheap path. With no per-language config, getting 60β130M tokens of Python/JS
from
stack-v3-trainmeans streaming whole-repo rows, opening a nestedfilesstruct, and filtering on a row field β for a source that concentrates in XML/HTML/JSON anyway. Code is not mandatory (Β§3.4 makes proportions our call), and at 100M params / 1B tokens the marginal value is genuinely unclear. Action: a dedicated mini-probe before committing β measure rows-sampled-per-Python-repo and the nested schema β and hold code at β€60M in the plan until that returns. If it is awkward, drop it; that is a defensible Β§3.4 choice, not a failure. eng_Latnandopen-web-mathare not actually English-only (8.5 % and 7.5 % of sampled rows fail a crude English test; the finepdfs sample included Cyrillic inline in a single "English" document). Β§3.4 is English-only, so a language filter is now a required pipeline stage, not a courtesy β applied per row on our side, on top of whatever the upstream config claims.finepdfscarriesfull_doc_lid/per_page_languagescolumns to do this cheaply; the general fallback islanguage_scorewhere present and a fastText-style gate where not.
9.4 Revised totals
Confirmed-and-usable now: 570M target tokens across 13 resolved configs, β4.5 GB, at measured yields
of roughly 909k documents β plus fineweb-edu/finewiki/wikipedia-monthly/finemath recoverable
with the corrected config names (Β§9.1), which restores ~570M of the planned 1,310M. Net effect on the
plan: the shape survives, the line items do not. Β§2's per-source token targets stay as targets;
Β§2's config strings and two column names are superseded by this section.
Directly measured consequence for the build: 909k documents for 570M tokens β 627 tokens/document average, so the mix is ~1.6β2.1M documents for 1.0β1.3B tokens. That is the number the dedup design has to be sized against β under 2.1M documents, a MinHash signature table comfortably fits the 30 GiB RAM, which materially de-risks the one stage Β§8 called the real cost.
10. Builder implemented, and one blocking flaw found and fixed
code/build/build_mix.py (mirrored in Cion-lab/ounce100m-code) is the Phase 2 builder. It stages each
source separately (stream β normalise β gates β dedup β tokenise) and then merges the staged shards
into the final mixed shards by largest-deficit round-robin over documents.
E-016, and why the smoke test was mandatory. The first version consumed one source to its target and
only flushed a shard when the buffer filled. Since fineweb-edu's target alone is ~37 shards of 8M tokens,
the first third of the mix would have been nothing but web text, then textbooks, then math β a curriculum
drift across the run that no LR schedule can repair, and one the builder itself reported as success
(checksums, headers, token counts and cursor probes were all clean). The verifier caught it as
distinct_sources_per_shard_min: 1. The shuffle that was supposed to fix ordering only permuted documents
inside a single-source buffer. After the two-stage rewrite, the same smoke run reports
min = max = median = 4 sources in every shard, 12,001,386 tokens recounted from the bytes,
0 out-of-range ids, 40/40 cursor probes and 40/40 document-boundary reconstructions correct.
Shard format (shared by staged and final, so document boundaries survive to the trainer):
header : <IIII = vocab, n_docs, n_tokens, flags 16 B
offsets : uint32 Γ (n_docs + 1), cumulative, [0] = 0
ids : uint16 Γ n_tokens (vocab 49,152 < 65,536, so uint16 is exact and halves the bytes)
doc i = ids[offsets[i] : offsets[i+1]]
Target 8M tokens/shard (16 MB, ~150 shards, ~2.5 GB for the mix). The resume cursor is
(shard_index, token_offset) β one integer pair, which is what makes Β§3.1's "same data, same position"
checkable and what keeps Trainer's replaying skip_first_batches cheap on a map-style read.
Interruption safety, which was not optional. /kaggle/working is destroyed when a session ends, and a
full staging pass is hours of streaming. So each source is published to Cion-lab/ounce100m-mix-stage
(the stage repo, distinct from the final ounce100m-mix-v1) the moment it completes, with a
record.json holding its shard list, token count and drop statistics. A new session rebuilds its stage
bookkeeping from the Hub and skips every source already published; only the source in flight at the moment
of the kill is re-staged. code/build/smoke_hub_resume.py rehearses this by wiping the local tree between
stage and merge, because "merge-only with an empty state and no restore path" is a failure mode that exits
with code 0 while emitting an empty mix.
Revised raw total: 1.23 B tokens = 23 % headroom over the 1.0 B trained target (raised fineweb-edu to 340M, finepdfs-edu to 180M, finemath to 160M, because Β§9's target edits had quietly pulled the plan from 1.31 B down to ~1.13 B β about 11 % headroom, which is not headroom once a 12β18 % dedup loss is applied). 16 sources, no code source yet (Β§9.3).
Gates the build must satisfy before Gate 2 is claimed, as assertions in verify_mix.py rather than
prose: every shard sha256 matches the manifest; no trailing bytes; offsets well-formed; zero out-of-range
ids; total tokens recounted from bytes equals the manifest; and β₯ half the distinct sources present in
every shard. The last one exists because the first build passed everything else.
11. Interruption safety proven, and the build that this authorises
A CPU session's /kaggle/working is destroyed when it ends. Everything below was measured, not assumed
(dodosoomro/ounce100m-p2-cold-resume-rehearsal v6, 17:35Z, 55 s wall-clock, pinned REV 428c5234).
| Step | Observed |
|---|---|
| 1. stage two sources, publish each | 2,400,867 tokens staged; record.json + shard-0000.bin per source on the Hub |
2. wipe the local tree (root_exists: false, stage_dir_exists: false) |
faithful simulation of a session kill, not a soft restart |
3. --merge-only with the Hub as the only state |
hub: restored 2 staged source(s) (2,400,867 tokens) β merged shard 2,352,747 tok / 5,815 docs, 2/2 sources in it, val shard 48,120 tok |
4. verify_mix.py over the restored mix |
PASS: true, sha_ok 1/1, out_of_range_ids: 0, trailing_bytes: 0, cursor probes 40/40 boundaries reconstructed, 0 mismatches, recount_matches_manifest: true |
| 5. verdict + self-cleanup | COLD_RESUME_PASSED=True, cleanup=deleted, throwaway repo confirmed gone from the namespace |
v4 of the same rehearsal had reported false over an identically-working pipeline: a stray
print("shards=", β¦) inside the BUILD_JSON block made the parent's json.loads fail and a helper fold
that into None (memory/ERRORS.md E-019, and the report-shape recurrence E-023). The fix was to
the harness, plus a (obj, why_null) return and a manifest.json fallback so a verdict can never again
be silently null.
The build now running (dodosoomro/ounce100m-p2-full-mix-build, 17:38Z, CPU, sessionTimeoutSeconds 43200, pinned REV 98e49560): build_mix.py --root /kaggle/working/mixroot --hub-repo Cion-lab/ounce100m-mix-stage with defaults β stage cap 1.31 B tokens, per-source targets summing to
1.23 B, 8 M tokens/mixed shard, VAL_STRIDE 50 holdout, shuffle seed 20260919. Each source publishes as
it completes, so an interruption costs at most one source's streaming time, and a re-launch restores the
finished ones from the Hub (step 3 above is that exact path).
Audit semantics, hardened the same way. audit_contamination.py compares 13-token windows against
{train, validation, dev} of the eight tasks β never test β and emits counts only. It now also reports
shards_missing, and its verdict is composite: AUDIT_PASSED requires every manifest shard present and
read, > 0 documents and grams measured, β₯ 6 readable reference sets, zero reference errors, and zero
overlap. Exit 6 means overlap found (a real finding); exit 7 means the audit did not measure (a bug).
A zero-overlap report over shards that were not on disk is no longer expressible.
Publishing ships the tokenizer. publish_mix.py copies tokenizer.json into the dataset root,
re-hashes it and aborts with exit 8 unless it matches the sha256 recorded in the manifest β so the
published ids and the published tokenizer are provably the pair that produced them (Β§3.12).
12. Gate 2 β the record, including where it failed and what that failure was worth
Merge and verify: passed, twice, reproducibly. From the 15 staged sources
(Cion-lab/ounce100m-mix-stage, 1,150,015,841 tokens staged):
| field | value |
|---|---|
| merged shards | 141 (each drawn from all 15 sources) |
| train tokens | 1,127,163,952 |
| held-out validation | 22,851,889 in 12 shards (every 50th merged document) |
| documents | 1,078,807 |
| per-source share | 75,144,263 tokens each, equal by construction |
| merge time | 16.4 s (restore from the Hub included: 123.8 s) |
verify_mix |
PASS: true β sha256 141/141, headers 141/141, 0 bad offsets, 0 out-of-range ids, 0 trailing bytes, recount matches manifest, 40/40 cursor probes, distinct_sources_per_shard_min = 15 (floor was 7) |
finewiki__en |
never staged. Gate 2's own diagnostic showed the columns present and first_row_ok: true with a truncated ImportError in the loading script. 15 sources at 1.127 B is above the 1.0 B target with 12.7 % headroom, so this is a composition note, not a blocker |
Then the audit refused to publish, and it was right to. audit_contamination exited 7:
| task | 13-gram hits | reference size (verified against the dataset server) |
|---|---|---|
| mmlu | 904,288 | dev 285 + validation 1,531 items β 149k grams |
| gsm8k | 185,600 | train 7,473 items = 360,656 grams |
| hellaswag | 1,852 | train 39,905 + validation 10,042 |
| arc_easy | 358 | train 2,251 + validation 570 |
| arc_challenge | 214 | train 1,119 + validation 299 |
| piqa | not measured | the loader yielded 0 of 16,113 train rows without raising |
| winogrande | (not in the visible excerpt) | train 40,398 + validation 1,267 |
| truthfulqa | (repo name was wrong) | validation 817, multiple_choice |
expected_false_positives for the whole scan was β0.0001, so these are matches, not collisions. The
first reading I wrote in STATE.md β "GSM8K's whole train split is in the mix" β was an over-claim: hits
are counted per mix gram, so a small number of frequently repeated reference grams can produce a large
count, and 185,600 hits against 360,656 distinct GSM8K grams means "up to half of its gram vocabulary
occurs somewhere in 1.1 B tokens". The columns that settle severity are mix_documents_majority_hit
(documents whose grams are more than half matched β verbatim inclusion) and source_hit_document_fraction
(which source, and how much of it) β and both are only computable on the fixed audit, which is what Gate
2c/2d exist to produce.
Why upstream did not catch it, verified rather than assumed: FineMath's card states it removed 13-gram
overlap against the test sets of GSM8k, MATH, MMLU and ARC; Cosmopedia's describes a 10-gram +
difflib pass against test benchmarks; OpenWebMath documents only SimHash self-deduplication with no
benchmark claim. validation/dev/train overlap was therefore never in scope for any of them β and
validation is the split lm-eval scores HellaSwag, PIQA, Winogrande and TruthfulQA on
(docs/05-eval-plan.md Β§7).
Gate 2 remained unsigned at the end of this stage. Cion-lab/ounce100m-mix-v1 did not exist yet. The
path to it: attribution β
drop the offending documents by mask (or a source if attribution says it is broadly soaked) β re-merge
with --exclude-dir β re-audit until AUDIT_PASSED with all 8 tasks covered β publish_mix, which now
refuses to ship unless the audit passed, covered all 8, scanned the val shards, and its mix_bytes_sha256
matches the exact manifest being published. E-030 explains why the first filtered re-merge could not
work and what was changed.
12.4 Resolution β the filtered mix, measured (Gates 2bβ2f, 2026-09-19)
Attribution first, always. Gate 2b/2c scanned the staged shards per source, which is the only way to
say where contamination lives rather than how much the merged mix contains. Of 1,955,832 staged documents,
10,425 (0.533 %) carried at least one matched 13-gram and 0 were majority-matched β shared spans,
not wholesale inclusion. The count concentrated in five sources: finemath-4plus 2,688 documents, fineweb-edu
1,043, cosmopedia auto_math_text 568, finepdfs-edu 420, open-web-math 339, with the remaining ~5,367
spread over ten more. Per-source hit counts summed to ~50,700 while the merged scan reported 2,096,540 for
essentially the same gram volume; the discrepancy is unreconciled and recorded as E-032c rather than
smoothed over. It changes no decision: the exclusion is per-document and the post-filter overlap is zero
either way.
The exclusion, and what it cost. audit_contamination --write-filter emitted one .u8 mask per source
(1 byte per staged document, ~700 KB total) marking documents that contain a matched window. Re-merging with
--exclude-dir dropped 5,878 documents = 17,366,967 tokens (1.5 % of the mix) and the re-audit of the
result gave overlap_total = 0 with 8/8 benchmark tasks covered, the held-out validation shards
scanned clean, and MEASURED: true. Final mix: 1,109,714,831 training tokens in 141 shards plus
22,851,889 held-out validation tokens β still 11.0 % above the 1.0 B target, so the pre-registered
0.9 B fallback (D-012) did not fire and the contamination fix cost nothing that the token budget had to
pay for.
Reference coverage is now a precondition, not an assumption. Gate 2b found that piqa's train and
validation splits returned 0 rows from the datasets library inside the Kaggle image while the server
reports 16,113 + 1,838 β the first audit had scored that as "clean". The audit now fetches each split's
server-side row count (/size), refuses any split it read to under 90 % of it, retries through the HTTP
/rows endpoint, and marks the whole report MEASURED: false if a loader still falls short. TruthfulQA's
answer text also sits under keys (mc1_targets, mc2_targets, abstractions) the first version never
read. All reference splits used are train/validation/dev; no test split was opened (Β§3.3).
Why publication took three attempts (E-033). Gates 2e and 2f both built and audited the mix correctly
and then died in the uploader: publish_mix sent one commit per file, and the Hub began rejecting every
request β including 30-byte files β after ~139 commits. That is a commit-rate ceiling, not a network fault,
and the "resumable" re-send was itself broken because RepoFile.lfs is a dict and getattr(ent.lfs, "sha256", None) is therefore None for every shard, i.e. nothing looked published. The uploader now
reads lfs["sha256"] or lfs["oid"], sends the remaining files as one upload_folder commit, and after
the commit re-lists the repo and re-hashes, exiting 11 if a path is absent and 12 if a remote digest
disagrees with the local bytes. Cion-lab/ounce100m-mix-v1 held 142 of 173 files when Gate 2g re-ran the
same pipeline with it.
Gate 2 is signed (dodosoomro/ounce100m-gate-2g-publish-mix-v1, complete 2026-09-20T00:31Z, CPU, zero
GPU quota). The rebuilt mix reproduced Gate 2e's token total exactly β 1,109,714,831 β which is the
determinism claim in Β§11 tested a second time on a different session with the same inputs. audit.json on
the Hub reports overlap_total = 0, mix_documents_with_any_hit = 0, majority_hit_docs = 0,
expected_false_positives = 0.0, tasks_covered = all eight, val_hits = 0 across 21,898 held-out
documents, and AUDIT_PASSED, MEASURED, COVERED_ALL_SHARDS all true, bound to these exact bytes by
mix_bytes_sha256 = 732eec7f19e77f36β¦. verify_mix exited 0. The batched uploader landed all 174 files
(139 train shards, 12 val shards, 16 filter/ objects, manifest.json, audit.json, tokenizer.json,
README.md, ATTRIBUTION.md, build_state.json, .gitattributes), and four train plus two validation
shards were then read back from the anonymous resolve endpoint and matched the manifest's sha256 β so the
mix is public and serving, not merely listed (Β§3.13).
One editorial note on this section: Β§12.2's first draft treated overlap_total as a count of contaminated
documents. It is a count of matched grams, which is why a 41Γ difference between the per-source and
merged hit totals could coexist with a 1.3 % agreement on document counts (E-032c). The action taken β
drop the documents, re-audit to zero β was unaffected, but the number in the record was read wrongly for
two gates, and the fix was to write the density columns (mix_documents_with_any_hit,
mix_documents_majority_hit) and quote those when judging severity.