nohi191212 commited on
Commit
5ce4e14
·
verified ·
1 Parent(s): 0a1670a

Point to current paper and 80 main-table router weights

Browse files
Files changed (1) hide show
  1. README.md +18 -53
README.md CHANGED
@@ -1,67 +1,32 @@
1
  ---
2
- language:
3
- - en
4
  tags:
5
- - computer-vision
6
- - model-routing
7
- - reproducibility
8
  ---
9
 
10
- # ModelCollaboration: five-seed results, router weights, and paper
11
 
12
- **Current release: 13 September 2026, paper revision 32.** The current method uses four-state cross-entropy training and the routing score **P(rescue) − P(harm)**. It was selected by task-balanced validation Score25 **mean − sample standard deviation** across seeds **2026–2030**. Current main results average all five runs; the full objective comparison contains **5 methods × 16 model pairs × 5 seeds = 400 router fits**. The 80 main-method routers are a subset of those 400.
13
 
14
- The current replay and experiment sources are shipped **in this Hugging Face repository**. The [original GitHub repository](https://github.com/nohi191212/ModelCollaboration) retains the earlier implementation and setup documentation; its original single-seed result tables and replay entry are not the current five-seed release.
15
 
16
- ## Current files
17
 
18
- | File or directory | Contents |
19
- |---|---|
20
- | `router_weights_five_seeds.tar.gz` | All 400 router checkpoints, run configurations, and training-set normalization parameters; extracts into `router_weights_five_seeds/runs/`. |
21
- | `experiment_data_five_seeds.tar.gz` | All 400 validation/test raw router-logit records, exact float32 routing scores, configurations and evaluation evidence, plus shared per-example task outcomes and sample identifiers for the 16 model pairs; extracts into `experiment_data/`. |
22
- | `scripts/replay_five_seed_results.py` | Current NumPy-only replay: recomputes scores, validation-selected operating points, curves, five-seed means/SDs, and objective selection from the released evidence. |
23
- | `paper_results/` | Revision-32 data, source tables, figures, and plotting sources. `data/results.json` is the current main-table/curve summary. `data/method_lock_five_seeds.json` records the final objective-selection rule and values. |
24
- | `main.pdf`, `appendix.pdf` | Current paper and full supplementary material, including the complete validation/test objective comparison. |
25
- | `reproduction/` | Training, evaluation, and five-seed selection sources with explicit local paths and safe import behavior. These sources require the original prepared feature caches for retraining; those caches are not part of the compact replay archive. |
26
- | `docs/` | [Five-seed protocol](docs/FIVE_SEED_PROTOCOL.md), [artifact and source guide](docs/ARTIFACTS.md), and [release changes](docs/RELEASE_NOTES.md). |
27
 
28
- ## Recompute the current results without running any models
 
 
 
29
 
30
- Download the current small files and the new evidence archive into one directory. The router weights are **not required** for this numerical replay.
31
 
32
- ```bash
33
- pip install huggingface_hub numpy
34
- hf download nohi191212/ModelCollaboration --local-dir modelcollaboration \
35
- --include "README.md" "docs/*" "scripts/replay_five_seed_results.py" \
36
- "experiment_data_five_seeds.tar.gz" "paper_results/*" "main.pdf" "appendix.pdf"
37
- cd modelcollaboration
38
- tar -xzf experiment_data_five_seeds.tar.gz
39
- python scripts/replay_five_seed_results.py --artifacts . --output-dir replayed
40
- ```
41
 
42
- The replay writes `replayed_main_five_seeds.csv`, `replayed_objectives_five_seeds.csv`, `replayed_per_seed.csv`, `replayed_same_checkpoint_controls.csv`, and `verification_five_seeds.json`. It does not execute the specialist, VLM, or router. It independently checks the score transformations from raw logits, then uses saved float32 scores to preserve the original threshold-boundary decisions when recomputing task metrics. ConstructionSite is evaluated by event macro-F1; the other tasks use their saved correctness outcomes. The [artifact guide](docs/ARTIFACTS.md) explains the precision check.
43
 
44
- The release verification passed for all **400 fits**, **160 same-checkpoint scoring controls**, **118,160 learned-routing curve points**, and **3,232 confidence-routing curve points**, including the current main table and five-seed means/sample SDs. Confidence scores and validity flags are included for independent replay of that baseline. See [`verification_five_seeds.json`](verification_five_seeds.json) and [`replayed/`](https://huggingface.co/nohi191212/ModelCollaboration/tree/main/replayed) for the verified outputs.
45
 
46
- Validation thresholds are transferred unchanged to test samples. A deployed threshold can decide for each arriving sample; it does not require access to all test inputs. Exact-budget test rankings are reported separately as batch diagnostics. See the [protocol](docs/FIVE_SEED_PROTOCOL.md) for this distinction and the documented earlier test inspection.
47
-
48
- ## Download the trained routers
49
-
50
- ```bash
51
- hf download nohi191212/ModelCollaboration router_weights_five_seeds.tar.gz --local-dir .
52
- tar -xzf router_weights_five_seeds.tar.gz
53
- ```
54
-
55
- Current main-method directories end in `four_difference_s2026` through `four_difference_s2030`. The other directories contain the four matched objective/scoring controls. A checkpoint must be used together with its configuration, training normalization, and corresponding frozen feature encoders. Per-run thresholds belong to that run; an average threshold is not a replacement for the five individual policies.
56
-
57
- ## What changed in the evidence
58
-
59
- At the validation-selected test operating points, the current five-seed means outperform confidence routing in **13/16** pairs, use fewer VLM calls in **11/16**, achieve both in **10/16**, and outperform both stand-alone models in **8/16**. These are comparisons of means, not a claim that every difference is statistically significant. Standard deviations describe router-training randomness with frozen encoders, specialist outputs, and VLM outputs.
60
-
61
- For the two cost examples in Fig. 3, InstanceVG/MiniCPM improves the metric by **6.15 percentage points**, with **70.8% fewer estimated FLOPs** and **68.6% less estimated serial time** than confidence routing. YOLO26x/Qwen improves by **3.90 points**, with reductions of **53.0%** and **52.5%**. Costs include the frozen routing encoders and routing head on every input. They reuse measured A800 component costs (64 inputs, batch one) and scale VLM cost by the actual call rate; they are not a newly measured full-system deployment trace. Savings are pair-dependent: the current InstanceVG/Qwen selected policy has a high call rate and does not retain the previous single-seed saving claim.
62
-
63
- ## Earlier archives and external resources
64
-
65
- `weights.tar`, `experiment_data.tar.gz`, `replayed_main.csv`, and the earlier verification records are **legacy single-seed artifacts**. They must not be used to reconstruct the current five-seed tables. The unchanged distilled encoders and trained construction specialists in `weights.tar` are still reusable: `encoders/8M/` and `specialists/` remain relevant, while its `main/runs/` contains the old 16 routers.
66
-
67
- Original images, upstream VLM weights, some specialist weights, and full prepared feature caches remain external. Obtain the original resources and preparation instructions from the [GitHub setup documentation](https://github.com/nohi191212/ModelCollaboration/tree/main/docs), then use the current scripts in `reproduction/` for the new objective study. This release supports compact numerical replay without large-model inference; it does not claim one-command reproduction from raw images. Original dataset/model licenses and access conditions continue to apply; this model card grants no new license to third-party assets.
 
1
  ---
2
+ license: mit
 
3
  tags:
4
+ - model-collaboration
5
+ - learning-to-defer
6
+ - vision-language
7
  ---
8
 
9
+ # Learned Routing for Specialist–VLM Collaboration
10
 
11
+ Artifacts for **When Should a Vision Specialist Defer? Budgeted Collaboration with Vision–Language Models**.
12
 
13
+ [Code and reproduction instructions](https://github.com/nohi191212/ModelCollaboration) · [Latest paper](releases/20260913_lr/main.pdf) · [Supplement](releases/20260913_lr/appendix.pdf)
14
 
15
+ ## Latest release
16
 
17
+ Use [releases/20260913_lr](https://huggingface.co/nohi191212/ModelCollaboration/tree/main/releases/20260913_lr).
 
 
 
 
 
 
 
 
18
 
19
+ - **80 main-table router checkpoints:** 16 model pairs × five seeds, in `full_budget_20260913_weights.tar.gz`.
20
+ - Experiment code, configurations, histories and per-example outputs, including the other objectives and scoring experiments. Only the main-table router weights are included in this release.
21
+ - Eight specialist input caches covering train, validation and test splits, with an input manifest.
22
+ - Paper sources and numerical summaries in the code repository; latest PDFs and result data here.
23
 
24
+ To reconstruct Tables 1 and 3 without training, download `paper_replay_runs.tar.gz` and `paired_outcomes.tar.gz`, then follow the [reproduction instructions](https://github.com/nohi191212/ModelCollaboration/blob/main/docs/REPRODUCE.md). The replay checks 80 main runs, 320 Table 3 runs and 19,392 curve points.
25
 
26
+ Main checkpoints use validation selection across 0–100% calls. Table 3 preserves its earlier checkpoint selection settings, detailed in the supplement.
 
 
 
 
 
 
 
 
27
 
28
+ Root-level archives are retained for earlier releases. The latest release does not add supplementary router checkpoints.
29
 
30
+ ## Data and licenses
31
 
32
+ Original code is MIT licensed. Upstream model and dataset licenses remain unchanged. This release provides derived experiment features, labels and outputs; obtain original images from their providers. ConstructionSite-derived data retain CC BY-NC 4.0 terms. See [model and dataset sources](https://github.com/nohi191212/ModelCollaboration/blob/main/docs/MODELS_AND_DATA.md).