ModelCollaboration / docs /ARTIFACTS.md
nohi191212's picture
Publish five-seed probability-difference results, 400 routers, and revision 32
466c2b6 verified
|
Raw History Blame Contribute Delete
8.08 kB

Artifact layout and experiment source guide

Minimal numerical replay

Download experiment_data_five_seeds.tar.gz, scripts/replay_five_seed_results.py, and paper_results/ from this repository. Extract the archive into the download root. The --artifacts argument points to that root, containing both experiment_data/ and paper_results/data/, not to experiment_data/ itself.

download-root/
  scripts/replay_five_seed_results.py
  experiment_data/
    runs/repeat_<expert>_<endpoint>_<method>_s<seed>/
      config.json
      paired_logits.npz
      replay_scores.npz
      completion.json
      test_evaluation.json
      ...
    pairs/shared_12_<expert>_<endpoint>/
      ... shared validation/test outcomes and sample identifiers ...
  paper_results/data/
    results.json
    method_lock_five_seeds.json
    per_seed.csv
    pair_mean_std.csv
    objective_mean_std.csv
    ...

paired_logits.npz stores the raw validation and test outputs before the score transformation. replay_scores.npz stores routing scores derived from those logits with the original PyTorch CPU float32 operations; keys combine split and ranking, such as validation_probability_difference. Producing these score arrays executes no model and retrains nothing. They preserve the original decisions at float32 threshold boundaries: numerically close softmax implementations can otherwise move a tied value across a threshold.

The NumPy-only replay independently checks every raw-logit score formula against these saved scores (maximum tolerance 8e-7 for cross-implementation float32 rounding), then uses the exact saved scores to recompute all selected policies, calls, and metrics (tolerance 1e-12 against released numerical evidence). The latter check is not relaxed by the score-formula tolerance. The shared pair data specify the task outcomes needed for accuracy or construction macro-F1. They also store the original correctness labels separately: an invalid construction answer can have the same TP/FP/FN counts as a correct negative answer, so those counts alone cannot reconstruct rescue/harm membership. Per-run configurations identify the method, seed, target, ranking score, architecture, and optimizer choices. The reference paper data and current replay verification jointly check the published summaries; they do not replace the per-example arrays.

Each shared val.npz/test.npz contains small, large, small_correct, large_correct, confidence, and confidence_valid. small/large contain correctness vectors for accuracy tasks or per-event TP/FP/FN arrays for ConstructionSite. The explicit correctness arrays have shape (N, 1) or (N, 4). Confidence ranking uses -confidence for valid specialist confidence and 1 otherwise, exactly as in the original audit. Both policies use validation-derived thresholds and are independently replayed on test.

Checkpoint use

The new weights archive extracts separately:

router_weights_five_seeds/
  runs/repeat_<expert>_<endpoint>_<method>_s<seed>/
    best.pt
    config.json
    hidden_normalization.npz
    confidence_normalization.npz

All 400 run names correspond to the evidence archive. The 80 four_difference runs are the current main method. Merge corresponding run directories into a new working directory if evaluating checkpoints with their saved validation logits: the weight archive provides checkpoints/normalization and the evidence archive provides validation_logits.npy and other records. Preserve released evidence as read-only source data.

The frozen 8M encoder and construction specialist checkpoints did not change. They remain in the legacy weights.tar archive under encoders/8M/ and specialists/. Its main/runs/ contains 16 earlier single-seed routers and must not be substituted for the new 80 main routers.

Training and evaluation source

The reproduction/ directory contains the executed experiment algorithms with local path arguments and main guards added for portability:

  • train_objective.py: trains one router from a released configuration using train/validation caches only.
  • evaluate_objective.py: reloads one checkpoint, checks its saved validation outputs, transfers validation thresholds to test, and evaluates alternate scores on the same four-state checkpoint.
  • select_five_seeds.py: reads 400 validation completion records and selects the global method by mean minus sample SD. It does not read test records.
  • run_study.py: optional explicit launcher for the 400 configurations; it trains all runs, performs final validation-only selection, then evaluates test data. It omits the old seed-42 pilot from the current entry point.
  • routing/train_router_formal.py: shared architecture, target construction, feature loader, and metrics. Its older standalone main() is retained as historical source; use the current entry points above.

No training runs on module import. The release update checks syntax and argument help; it does not rerun the 400 fits. The launcher has explicit resource/path arguments and requires a new output directory. It is an optional retraining entry, not required by the NumPy-only replay.

Prepared features required for retraining

--root points to a prepared project with these original cache conventions:

prepared-project/
  outputs/router_full_20260910/
    data/<expert>/<train|val|test>/
      records.jsonl
      confidence.npy
      output.npy
      index.json
    features/8M/<task>/<train|val|test>/
      sample_ids.json
      image.npy
      text.npy
      valid.npy
  outputs/mini_siglip_strict_20260910_300ep/8M/config.json

Each hidden-state index.json points to matching states.npz shards. Paths inside these indices must point to the user's own prepared shards. The feature/sample-ID order must match the records; the loader checks that boundary. These full train/validation/test feature caches are not in the compact evidence archive. Original upstream model/dataset downloads and preprocessing instructions remain in the original setup documentation; apply the current supplementary split definitions and task prompts. Some specialist weights remain external. This release does not provide a one-command raw-image reconstruction of all caches.

The minimal replay download omits training source. Download it before retraining:

hf download nohi191212/ModelCollaboration --local-dir . --include "reproduction/*"

With those caches prepared, a single fit can be run explicitly:

python reproduction/train_objective.py --root /path/to/prepared-project \
  --config experiment_data/runs/repeat_cub_Qwen3.8_four_difference_s2026/config.json \
  --output /path/to/new-run --device cuda
python reproduction/evaluate_objective.py --root /path/to/prepared-project \
  --run /path/to/new-run --device cuda

The optional all-run launcher is:

python reproduction/run_study.py --root /path/to/prepared-project \
  --artifacts /path/to/download-root --output /path/to/new-study --gpus 0 1

This starts 400 fits and evaluations on the GPUs explicitly supplied. It is not part of the release verification command in the main README. Training requires PyTorch and NumPy in addition to the prepared features.

Historical provenance and licenses

experiment_data.tar.gz, weights.tar, and replayed_main.csv document the earlier seed-42 release. The original pilot selected probability-ratio scoring; the final five-seed study selects probability difference. The protocol records the amended rule and earlier test access. Do not treat the new evidence as a newly blinded test set.

Original images and upstream weights are not republished in the compact data archive. All upstream licenses and access conditions remain applicable. The unchanged frozen encoders and construction weights are stored in the original weights archive; retaining that archive avoids duplicating them in the new router-only package.