Point to current paper and 80 main-table router weights
Browse files
README.md
CHANGED
|
@@ -1,67 +1,32 @@
|
|
| 1 |
---
|
| 2 |
-
|
| 3 |
-
- en
|
| 4 |
tags:
|
| 5 |
-
-
|
| 6 |
-
-
|
| 7 |
-
-
|
| 8 |
---
|
| 9 |
|
| 10 |
-
#
|
| 11 |
|
| 12 |
-
|
| 13 |
|
| 14 |
-
|
| 15 |
|
| 16 |
-
##
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|---|---|
|
| 20 |
-
| `router_weights_five_seeds.tar.gz` | All 400 router checkpoints, run configurations, and training-set normalization parameters; extracts into `router_weights_five_seeds/runs/`. |
|
| 21 |
-
| `experiment_data_five_seeds.tar.gz` | All 400 validation/test raw router-logit records, exact float32 routing scores, configurations and evaluation evidence, plus shared per-example task outcomes and sample identifiers for the 16 model pairs; extracts into `experiment_data/`. |
|
| 22 |
-
| `scripts/replay_five_seed_results.py` | Current NumPy-only replay: recomputes scores, validation-selected operating points, curves, five-seed means/SDs, and objective selection from the released evidence. |
|
| 23 |
-
| `paper_results/` | Revision-32 data, source tables, figures, and plotting sources. `data/results.json` is the current main-table/curve summary. `data/method_lock_five_seeds.json` records the final objective-selection rule and values. |
|
| 24 |
-
| `main.pdf`, `appendix.pdf` | Current paper and full supplementary material, including the complete validation/test objective comparison. |
|
| 25 |
-
| `reproduction/` | Training, evaluation, and five-seed selection sources with explicit local paths and safe import behavior. These sources require the original prepared feature caches for retraining; those caches are not part of the compact replay archive. |
|
| 26 |
-
| `docs/` | [Five-seed protocol](docs/FIVE_SEED_PROTOCOL.md), [artifact and source guide](docs/ARTIFACTS.md), and [release changes](docs/RELEASE_NOTES.md). |
|
| 27 |
|
| 28 |
-
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
-
|
| 33 |
-
pip install huggingface_hub numpy
|
| 34 |
-
hf download nohi191212/ModelCollaboration --local-dir modelcollaboration \
|
| 35 |
-
--include "README.md" "docs/*" "scripts/replay_five_seed_results.py" \
|
| 36 |
-
"experiment_data_five_seeds.tar.gz" "paper_results/*" "main.pdf" "appendix.pdf"
|
| 37 |
-
cd modelcollaboration
|
| 38 |
-
tar -xzf experiment_data_five_seeds.tar.gz
|
| 39 |
-
python scripts/replay_five_seed_results.py --artifacts . --output-dir replayed
|
| 40 |
-
```
|
| 41 |
|
| 42 |
-
|
| 43 |
|
| 44 |
-
|
| 45 |
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
## Download the trained routers
|
| 49 |
-
|
| 50 |
-
```bash
|
| 51 |
-
hf download nohi191212/ModelCollaboration router_weights_five_seeds.tar.gz --local-dir .
|
| 52 |
-
tar -xzf router_weights_five_seeds.tar.gz
|
| 53 |
-
```
|
| 54 |
-
|
| 55 |
-
Current main-method directories end in `four_difference_s2026` through `four_difference_s2030`. The other directories contain the four matched objective/scoring controls. A checkpoint must be used together with its configuration, training normalization, and corresponding frozen feature encoders. Per-run thresholds belong to that run; an average threshold is not a replacement for the five individual policies.
|
| 56 |
-
|
| 57 |
-
## What changed in the evidence
|
| 58 |
-
|
| 59 |
-
At the validation-selected test operating points, the current five-seed means outperform confidence routing in **13/16** pairs, use fewer VLM calls in **11/16**, achieve both in **10/16**, and outperform both stand-alone models in **8/16**. These are comparisons of means, not a claim that every difference is statistically significant. Standard deviations describe router-training randomness with frozen encoders, specialist outputs, and VLM outputs.
|
| 60 |
-
|
| 61 |
-
For the two cost examples in Fig. 3, InstanceVG/MiniCPM improves the metric by **6.15 percentage points**, with **70.8% fewer estimated FLOPs** and **68.6% less estimated serial time** than confidence routing. YOLO26x/Qwen improves by **3.90 points**, with reductions of **53.0%** and **52.5%**. Costs include the frozen routing encoders and routing head on every input. They reuse measured A800 component costs (64 inputs, batch one) and scale VLM cost by the actual call rate; they are not a newly measured full-system deployment trace. Savings are pair-dependent: the current InstanceVG/Qwen selected policy has a high call rate and does not retain the previous single-seed saving claim.
|
| 62 |
-
|
| 63 |
-
## Earlier archives and external resources
|
| 64 |
-
|
| 65 |
-
`weights.tar`, `experiment_data.tar.gz`, `replayed_main.csv`, and the earlier verification records are **legacy single-seed artifacts**. They must not be used to reconstruct the current five-seed tables. The unchanged distilled encoders and trained construction specialists in `weights.tar` are still reusable: `encoders/8M/` and `specialists/` remain relevant, while its `main/runs/` contains the old 16 routers.
|
| 66 |
-
|
| 67 |
-
Original images, upstream VLM weights, some specialist weights, and full prepared feature caches remain external. Obtain the original resources and preparation instructions from the [GitHub setup documentation](https://github.com/nohi191212/ModelCollaboration/tree/main/docs), then use the current scripts in `reproduction/` for the new objective study. This release supports compact numerical replay without large-model inference; it does not claim one-command reproduction from raw images. Original dataset/model licenses and access conditions continue to apply; this model card grants no new license to third-party assets.
|
|
|
|
| 1 |
---
|
| 2 |
+
license: mit
|
|
|
|
| 3 |
tags:
|
| 4 |
+
- model-collaboration
|
| 5 |
+
- learning-to-defer
|
| 6 |
+
- vision-language
|
| 7 |
---
|
| 8 |
|
| 9 |
+
# Learned Routing for Specialist–VLM Collaboration
|
| 10 |
|
| 11 |
+
Artifacts for **When Should a Vision Specialist Defer? Budgeted Collaboration with Vision–Language Models**.
|
| 12 |
|
| 13 |
+
[Code and reproduction instructions](https://github.com/nohi191212/ModelCollaboration) · [Latest paper](releases/20260913_lr/main.pdf) · [Supplement](releases/20260913_lr/appendix.pdf)
|
| 14 |
|
| 15 |
+
## Latest release
|
| 16 |
|
| 17 |
+
Use [releases/20260913_lr](https://huggingface.co/nohi191212/ModelCollaboration/tree/main/releases/20260913_lr).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 18 |
|
| 19 |
+
- **80 main-table router checkpoints:** 16 model pairs × five seeds, in `full_budget_20260913_weights.tar.gz`.
|
| 20 |
+
- Experiment code, configurations, histories and per-example outputs, including the other objectives and scoring experiments. Only the main-table router weights are included in this release.
|
| 21 |
+
- Eight specialist input caches covering train, validation and test splits, with an input manifest.
|
| 22 |
+
- Paper sources and numerical summaries in the code repository; latest PDFs and result data here.
|
| 23 |
|
| 24 |
+
To reconstruct Tables 1 and 3 without training, download `paper_replay_runs.tar.gz` and `paired_outcomes.tar.gz`, then follow the [reproduction instructions](https://github.com/nohi191212/ModelCollaboration/blob/main/docs/REPRODUCE.md). The replay checks 80 main runs, 320 Table 3 runs and 19,392 curve points.
|
| 25 |
|
| 26 |
+
Main checkpoints use validation selection across 0–100% calls. Table 3 preserves its earlier checkpoint selection settings, detailed in the supplement.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
+
Root-level archives are retained for earlier releases. The latest release does not add supplementary router checkpoints.
|
| 29 |
|
| 30 |
+
## Data and licenses
|
| 31 |
|
| 32 |
+
Original code is MIT licensed. Upstream model and dataset licenses remain unchanged. This release provides derived experiment features, labels and outputs; obtain original images from their providers. ConstructionSite-derived data retain CC BY-NC 4.0 terms. See [model and dataset sources](https://github.com/nohi191212/ModelCollaboration/blob/main/docs/MODELS_AND_DATA.md).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|