|
Download docs/05-eval-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 11.1 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/05-eval-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/05-eval-plan.md
-
curl -L -o 05-eval-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/05-eval-plan.md
11.1 kB
| > **STATUS 2026-09-21T14:30Z: NOT EXECUTED.** The owner stopped Phase 6 (D-020) before any task | |
| > produced a number, so no score exists anywhere in this project. Everything below stays as the frozen, | |
| > pre-registered protocol — it is a specification for a future run, and must not be read as a result. | |
| # 05 — Evaluation plan, pinned **before** any score exists | |
| > §5 Phase 6 says "pin the harness and per-task settings". The reason to write this file now, while no | |
| > model has been trained, is §3.3: shot counts, prompts and parsers become a scoring lever the moment | |
| > results exist. Everything here is a frozen choice with a reason, not a default I will tune later. | |
| ## 1. Harness pin (verified against PyPI today, not remembered) | |
| | Item | Pin | Why | | |
| |---|---|---| | |
| | Harness | `lm-eval` **0.4.13** (EleutherAI `lm-evaluation-harness`) | the reference implementation behind the published numbers of every ~100M-1B base model used as a target anchor in `01-plan.md` §7; and the only one of the three candidate stacks with canonical prompts for all eight tasks (incl. PIQA and ARC-Easy) | | |
| | Alternate, as a cross-check only | `lighteval` 0.13.0 | if the two disagree, the discrepancy is reported, not averaged | | |
| | Backend | `model=hf` (transformers), **not** vLLM | vLLM's support for Turing (sm 7.5) at ≥0.18 is unverified on this hardware, and p0c already showed what an unverified stack assumption costs. A 106M model needs no serving optimisation: the whole eval set is minutes of fp16 forward passes on one card | | |
| | Devices | **one** T4, no data parallelism | determinism beats speed here; two ranks add a gather path that can reorder few-shot examples | | |
| | dtype | `float16` | matches training precision's storage; bf16 does not exist on sm 7.5 | | |
| | Loading | `AutoModelForCausalLM` + `AutoTokenizer`, no `--trust-remote-code` | §8 forbids custom code; if it needs trust-remote-code the export is wrong and Phase 5 has failed first | | |
| | Chat template | **none** | base model, §3.9. A chat template would be an unearned capability | | |
| | Model source | the public Hub repo, fresh instance, empty cache | the same clean-room condition as Gate 5 | | |
| Command template (the exact string goes into the model repo with the results): | |
| ```bash | |
| lm_eval --model hf \ | |
| --model_args pretrained=Cion-lab/ounce100m-<n>m,dtype=float16,trust_remote_code=False \ | |
| --tasks <task-ids> --batch_size 8 --seed 42 --verbosity INFO \ | |
| --output_path results/ --log_samples | |
| ``` | |
| ## 2. Tasks, metrics and shot policy | |
| Eight tasks, one row each, metric names as the harness reports them: | |
| | Task | lm-eval task id (confirmed at run time through `TaskManager().task_index`; `--tasks list` is | |
| not a listing command in 0.4.13 — E-048, measured on free CPU 2026-09-20 12:23Z) | Metric | Shots | | |
| |---|---|---|---| | |
| | ARC-Challenge | `arc_challenge` | `acc` | task default (25) | | |
| | ARC-Easy | `arc_easy` | `acc` | task default (25) | | |
| | HellaSwag | `hellaswag` | `acc` | task default | | |
| | MMLU | `mmlu` | `acc` (macro over subjects) | task default | | |
| | TruthfulQA | `truthfulqa_mc2` (mc1 reported alongside) | `acc,none` | task default (0) | | |
| | Winogrande | `winogrande` | `acc` | task default | | |
| | PIQA | `piqa` | `acc` | task default | | |
| | GSM8K | `gsm8k` (`main`, `flexible-extract`) | `exact_match,flexible-extract` | task default | | |
| **Shot policy, frozen.** PRIMARY results use each task's **harness default `num_fewshot`** with no | |
| `--num_fewshot` override anywhere in the run — the point of using the file in the repo rather than a | |
| number I typed is that nobody chose it per task. A SECONDARY uniform 5-shot column is reported for every | |
| task in the same run. Both columns ship; the headline is PRIMARY. Deciding this now means "re-run with | |
| more shots until MMLU moves" is not available as an option later. | |
| Seeds: `--seed 42` fixed for the whole table, one invocation per task, all eight in a single job so a | |
| partial run cannot be presented as a complete one. Incremental result files are pushed after each task | |
| completes (§5 Phase 6: an interruption must not lose finished work). | |
| ## 3. Expected outcome, and how it will be reported honestly | |
| The bands in `01-plan.md` §7 / D-005 were set from base models of comparable size trained on ~300 B | |
| tokens; this model sees ~1 B. Several tasks are therefore expected to sit **at or near chance**: | |
| MMLU (~25 %), WinoGrande (~50 %), ARC-C (~25 %), GSM8K (~0 %), TruthfulQA mc2 low single digits. That is | |
| the pre-registered prediction, and a table that lands there is a *success*, not a disappointment to be | |
| engineered around. Anything materially above band gets a contamination re-check before it is published, | |
| not a celebration — an unexpectedly high score is more likely to be a leak than a breakthrough. | |
| Reporting rules, also frozen now: | |
| 1. All eight, always, with metric name + shot count + the command above. | |
| 2. No subset, no "best of two runs", no per-task parser swap. | |
| 3. Chance level stated next to every score so a reader can see the signal, if any. | |
| 4. `docs/05-eval-report.md` records deviations from this file verbatim, including any that were forced. | |
| ## 4. Contamination posture, restated at the point of use | |
| Test splits are untouched until this phase (§3.3) and are read only by the harness, never by me. The mix | |
| was audited mechanically against train/validation/dev material of these same tasks (`build/audit_contamination.py`, | |
| 13-token windows, counts only, no item text ever surfaced). The final report must carry that audit's | |
| numbers next to these scores, because the two together are the only defensible statement about the | |
| results. | |
| ## 5. Sources checked today | |
| - `lm-eval` 0.4.13 and `lighteval` 0.13.0 read from the PyPI JSON API at 18:17Z on 2026-09-19, not from | |
| memory (§3.6). `lm-eval`'s `hf` extra declares `transformers>=4.1`, `accelerate>=0.26.0`; the Kaggle | |
| image carries transformers **5.0.0** (measured in preflight `pargs`), so compatibility is asserted at | |
| run time by a `--limit 5` smoke test on all eight tasks before the full table is produced. | |
| ## 6. Amendment after Gate 3 (D-011): the model's context is 1024, and that is reported | |
| The sequence length was changed from 2048 to 1024 on measured throughput (D-011), *before* the run was | |
| frozen. Consequence for this plan, recorded now so it cannot become an excuse later: | |
| - `arc_challenge` and `arc_easy` at their default 25 shots produce prompts longer than 1024 tokens. | |
| lm-eval truncates from the left for causal models. The score is still the harness's own metric on the | |
| harness's own configuration; it is simply a score for a 1024-context model. | |
| - No shot count is being changed in response. Every other task's prompt fits. | |
| - If ARC's number looks odd next to a 2048-context anchor, the explanation is the context window, and it | |
| belongs in the model card and `final-report.md`, not in a re-tuned prompt. | |
| Measured on free CPU 2026-09-20 (E-049, `ounce100m-p6-cli-probe`): the nine flags this file's command lines use all exist in `lm-eval` 0.4.13's CLI, and the metric *names* below are what the pinned YAMLs declare — but the results **key** is `<name>,<aggregation>` and the aggregation is not settled by reading configs. So each row's key is verified by the `--limit 5` smoke pass before the full columns run, and a mismatch there is fixed in the constant, never by choosing a different metric after seeing a score. | |
| ## 7. Which split each task is actually scored on — read out of the pinned config, not assumed | |
| Added 2026-09-19T20:25Z, before any score exists, while Gate 2's contamination finding was being | |
| diagnosed. Every line below was read from `lm_eval/tasks/**` at tag **v0.4.13** through the GitHub API | |
| (file names listed from the tag, then the YAML itself) — not from memory and not from a README. | |
| | task config | scored on | few-shot exemplars from | `num_fewshot` declared | auditable by me? | | |
| |---|---|---|---|---| | |
| | `winogrande/default.yaml` | **validation** (`test_split` absent) | train | none | **yes — and it is in my reference set** | | |
| | `hellaswag/hellaswag.yaml` | **validation** (`test_split: null`) | train | none | **yes — in my reference set** | | |
| | `piqa/piqa.yaml` | **validation** (`test_split: null`) | train | none | **yes — was unreadable in Gate 2, fixed for 2c** | | |
| | `truthfulqa/truthfulqa_mc1/mc2` | **validation** (`test_split: null`) | none (0-shot) | **0** | **yes — in my reference set** | | |
| | `arc/arc_easy.yaml`, `arc_challenge` | test | **validation** | none | partially: val is in my reference set, test is not | | |
| | `mmlu/default/_default_template_yaml` | test | `fewshot_split: dev` | none | partially: dev is in my reference set | | |
| | `gsm8k/gsm8k.yaml` | test | train | **5** | partially: train is in my reference set | | |
| Three consequences, recorded now so they cannot be argued later: | |
| 1. **Four of the eight tasks are scored on `validation`** — Winogrande, HellaSwag, PIQA and TruthfulQA. | |
| §3.3 lets me read validation but never test, so those four are the ones a contamination audit can | |
| *fully* cover, and they are also the ones where overlap directly inflates a headline number. ARC, | |
| MMLU and GSM8K are scored on `test`, which stays unread until Phase 6: for them the audit covers | |
| exemplar material, not the scoring items. That asymmetry is a limitation of the method, not a choice | |
| about which results to check. | |
| 2. **"Each task's own default shot count" (D-010) needs its definition pinned, and here it is:** the | |
| `num_fewshot` value in the task's own config at v0.4.13 if present, otherwise the harness default of | |
| **0**. On the record before results: that makes TruthfulQA 0-shot by its own declaration, GSM8K 5-shot | |
| by its own declaration, and the remaining six 0-shot in the headline column — which is exactly why the | |
| pre-registered uniform **5-shot secondary column** exists. The primary metric stays the harness's own | |
| (accuracy for the multiple-choice tasks, `exact_match` for GSM8K, `mc2` for TruthfulQA). | |
| 3. **The upstream decontamination that the mix sources document is against *test* sets.** FineMath's card | |
| states: "Following Qwen2.5-Math's approach, we removed samples with 13-gram overlaps against **test | |
| sets** from GSM8k, MATH, MMLU and ARC" (with logs published as `HuggingFaceTB/finemath_contamination_report`), | |
| and Cosmopedia's card describes a 10-gram + `difflib` pass "against the test benchmarks". OpenWebMath's | |
| card documents SimHash deduplication only, with **no** benchmark decontamination claim. So the sources | |
| were cleaned for the splits I am forbidden to read, and left untouched for the four splits that | |
| Winogrande, HellaSwag, PIQA and TruthfulQA are actually scored on. That is consistent with Gate 2's | |
| measured hits and is the reason the audit found something real rather than imaginary: an overlap on | |
| `validation` is precisely what their filters were not looking for. | |
| Sources: [FineMath card](https://huggingface.co/datasets/HuggingFaceTB/finemath), | |
| [Cosmopedia card](https://huggingface.co/datasets/HuggingFaceTB/cosmopedia), | |
| [OpenWebMath card](https://huggingface.co/datasets/open-web-math/open-web-math). | |