# Artifact layout and experiment source guide ## Minimal numerical replay Download `experiment_data_five_seeds.tar.gz`, `scripts/replay_five_seed_results.py`, and `paper_results/` from this repository. Extract the archive into the download root. The `--artifacts` argument points to that root, containing both `experiment_data/` and `paper_results/data/`, not to `experiment_data/` itself. ```text download-root/ scripts/replay_five_seed_results.py experiment_data/ runs/repeat____s/ config.json paired_logits.npz replay_scores.npz completion.json test_evaluation.json ... pairs/shared_12__/ ... shared validation/test outcomes and sample identifiers ... paper_results/data/ results.json method_lock_five_seeds.json per_seed.csv pair_mean_std.csv objective_mean_std.csv ... ``` `paired_logits.npz` stores the raw validation and test outputs before the score transformation. `replay_scores.npz` stores routing scores derived from those logits with the original PyTorch CPU float32 operations; keys combine split and ranking, such as `validation_probability_difference`. Producing these score arrays executes no model and retrains nothing. They preserve the original decisions at float32 threshold boundaries: numerically close softmax implementations can otherwise move a tied value across a threshold. The NumPy-only replay independently checks every raw-logit score formula against these saved scores (maximum tolerance `8e-7` for cross-implementation float32 rounding), then uses the exact saved scores to recompute all selected policies, calls, and metrics (tolerance `1e-12` against released numerical evidence). The latter check is not relaxed by the score-formula tolerance. The shared pair data specify the task outcomes needed for accuracy or construction macro-F1. They also store the original correctness labels separately: an invalid construction answer can have the same TP/FP/FN counts as a correct negative answer, so those counts alone cannot reconstruct rescue/harm membership. Per-run configurations identify the method, seed, target, ranking score, architecture, and optimizer choices. The reference paper data and current replay verification jointly check the published summaries; they do not replace the per-example arrays. Each shared `val.npz`/`test.npz` contains `small`, `large`, `small_correct`, `large_correct`, `confidence`, and `confidence_valid`. `small`/`large` contain correctness vectors for accuracy tasks or per-event TP/FP/FN arrays for ConstructionSite. The explicit correctness arrays have shape `(N, 1)` or `(N, 4)`. Confidence ranking uses `-confidence` for valid specialist confidence and `1` otherwise, exactly as in the original audit. Both policies use validation-derived thresholds and are independently replayed on test. ## Checkpoint use The new weights archive extracts separately: ```text router_weights_five_seeds/ runs/repeat____s/ best.pt config.json hidden_normalization.npz confidence_normalization.npz ``` All 400 run names correspond to the evidence archive. The 80 `four_difference` runs are the current main method. Merge corresponding run directories into a **new working directory** if evaluating checkpoints with their saved validation logits: the weight archive provides checkpoints/normalization and the evidence archive provides `validation_logits.npy` and other records. Preserve released evidence as read-only source data. The frozen 8M encoder and construction specialist checkpoints did not change. They remain in the legacy `weights.tar` archive under `encoders/8M/` and `specialists/`. Its `main/runs/` contains 16 earlier single-seed routers and must not be substituted for the new 80 main routers. ## Training and evaluation source The `reproduction/` directory contains the executed experiment algorithms with local path arguments and main guards added for portability: - `train_objective.py`: trains one router from a released configuration using train/validation caches only. - `evaluate_objective.py`: reloads one checkpoint, checks its saved validation outputs, transfers validation thresholds to test, and evaluates alternate scores on the same four-state checkpoint. - `select_five_seeds.py`: reads 400 validation completion records and selects the global method by mean minus sample SD. It does not read test records. - `run_study.py`: optional explicit launcher for the 400 configurations; it trains all runs, performs final validation-only selection, then evaluates test data. It omits the old seed-42 pilot from the current entry point. - `routing/train_router_formal.py`: shared architecture, target construction, feature loader, and metrics. Its older standalone `main()` is retained as historical source; use the current entry points above. No training runs on module import. The release update checks syntax and argument help; it does not rerun the 400 fits. The launcher has explicit resource/path arguments and requires a new output directory. It is an optional retraining entry, not required by the NumPy-only replay. ### Prepared features required for retraining `--root` points to a prepared project with these original cache conventions: ```text prepared-project/ outputs/router_full_20260910/ data/// records.jsonl confidence.npy output.npy index.json features/8M/// sample_ids.json image.npy text.npy valid.npy outputs/mini_siglip_strict_20260910_300ep/8M/config.json ``` Each hidden-state `index.json` points to matching `states.npz` shards. Paths inside these indices must point to the user's own prepared shards. The feature/sample-ID order must match the records; the loader checks that boundary. These full train/validation/test feature caches are not in the compact evidence archive. Original upstream model/dataset downloads and preprocessing instructions remain in the [original setup documentation](https://github.com/nohi191212/ModelCollaboration/tree/main/docs); apply the current supplementary split definitions and task prompts. Some specialist weights remain external. This release does not provide a one-command raw-image reconstruction of all caches. The minimal replay download omits training source. Download it before retraining: ```bash hf download nohi191212/ModelCollaboration --local-dir . --include "reproduction/*" ``` With those caches prepared, a single fit can be run explicitly: ```bash python reproduction/train_objective.py --root /path/to/prepared-project \ --config experiment_data/runs/repeat_cub_Qwen3.8_four_difference_s2026/config.json \ --output /path/to/new-run --device cuda python reproduction/evaluate_objective.py --root /path/to/prepared-project \ --run /path/to/new-run --device cuda ``` The optional all-run launcher is: ```bash python reproduction/run_study.py --root /path/to/prepared-project \ --artifacts /path/to/download-root --output /path/to/new-study --gpus 0 1 ``` This starts 400 fits and evaluations on the GPUs explicitly supplied. It is not part of the release verification command in the main README. Training requires PyTorch and NumPy in addition to the prepared features. ## Historical provenance and licenses `experiment_data.tar.gz`, `weights.tar`, and `replayed_main.csv` document the earlier seed-42 release. The original pilot selected probability-ratio scoring; the final five-seed study selects probability difference. The [protocol](FIVE_SEED_PROTOCOL.md) records the amended rule and earlier test access. Do not treat the new evidence as a newly blinded test set. Original images and upstream weights are not republished in the compact data archive. All upstream licenses and access conditions remain applicable. The unchanged frozen encoders and construction weights are stored in the original weights archive; retaining that archive avoids duplicating them in the new router-only package.