this repo used to show the intervention effect of different baseline values. theorically, if the visual tower can perfectly model the visual world, the small intervention should not change the final VLM output. this dir is used to check the overlap of failure cases between different baseline values. Therefore, I will run inference in multiple online model: 1. Qwen/Qwen2.5-VL-3B-Instruct 2. Qwen/Qwen2.5-VL-7B-Instruct 3. Qwen/Qwen3.5-4B 4. Qwen/Qwen3.5-4B-Base Datasets: VSR: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/VSR/V7_testing.json - baseline: 0 - baseline: 0.05 - baseline: 0.1 - baseline: 0.15 - baseline: 0.2 - baseline: 0.25 Youtube_self_depth_QA - Level 1: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level1/test.jsonl - baseline: 0 - baseline: 0.05 - baseline: 0.1 - baseline: 0.15 - baseline: 0.2 - baseline: 0.25 - Level 2: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level2/test.jsonl - baseline: 0 - baseline: 0.05 - baseline: 0.1 - baseline: 0.15 - baseline: 0.2 - baseline: 0.25 - Level 3: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level3/test.jsonl - baseline: 0 - baseline: 0.05 - baseline: 0.1 - baseline: 0.15 - baseline: 0.2 - baseline: 0.25 all of code with respect to load data will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/dataset all of inference code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/inference all of inference results will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/results - it will contain each model's performance in each dataset in each baseline - each sample's input and output will be saved in the corresponding json file all of analysis code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/analysis all runnable sh scripts will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/sh --- ## Implementation At baseline `b` the model sees **one** image: the scene rendered from the baseline-`b` viewpoint. `b = 0` is the untouched original view, i.e. the no-intervention control. The question and the reference answer are identical across baselines, so the per-sample answers line up directly and the failure sets can be compared. Baselines are keyed canonically as `0.00 0.05 0.10 0.15 0.20 0.25`. The on-disk directories are spelled inconsistently (`0.10` vs `0.2`, and the VSR annotation uses `0.20` for the correspondence but `0.2` for the image); all of that is normalised in `dataset/baselines.py`, so any spelling can be passed on the command line. | split | n | image at baseline b | answer space | |---|---|---|---| | VSR (`V7_testing.json`) | 2195 | `00_data/gen_output_VSR/original_images` (b=0) / `.../generate_image/` | True / False | | Youtube level 1 | 821 | `level1/visualizations//` | A B C D | | Youtube level 2 | 764 | `level2/visualizations//` | A B C D (option letter) | | Youtube level 3 | 795 | `level3/visualizations//` | A B C D | The Youtube visualisations are already rendered per baseline with the annotation points warped into that view, so the marked points track the scene across the sweep. Level 2 is a four-option single choice: the question embeds four candidate orderings as (A)-(D) and the answer is the option letter. Note that the image labels its four *points* A-D as well, so a model may answer either with the option letter ("C") or by writing out that option's ordering ("B > C > A > D"). Only the letter is the requested format, but the ordering is unambiguous, so it is mapped back through the sample's own option table rather than scored wrong; the two are told apart by whether the first standalone letter is followed by a ">". The loader detects the format from the records themselves, so the older annotations that asked for the ordering directly (24 candidates, including the `.ordering_bak` copies) still load and score correctly. ### Answering modes `--scoring generate` (default, instruction-tuned models) greedily decodes and parses the generation. `--scoring likelihood` (base models, which do not follow the answer format) picks `argmax_c sum_t log P(c_t | prompt, c_` block at the generation prompt. Thinking is therefore **off** by default (`--enable_thinking` to turn it on); it is rejected outright under `--scoring likelihood`, where it would score the candidates as the opening words of a reasoning trace rather than as the answer. A generation that never reaches an answer normalises to something outside the candidate set. Those count as wrong, but are also tallied as `num_unparsed` / `unparsed_rate` in `metrics.json`, so a high parse-failure rate is visible instead of being read as low accuracy. If a model's unparsed rate is high, re-run it with `--scoring likelihood`. ### Layout ``` dataset/ baselines.py canonical baseline naming datasets.py per-baseline sample loading, prompts, answer parsing inference/ run_inference.py one model x one split x one baseline analysis/ analyze_baseline_overlap.py result/ //baseline_/{predictions.jsonl,metrics.json,logs/} sh/ run_all.sh the whole grid (4 models x 4 splits x 6 baselines) run_sweep.sh one model x one split, all baselines, sharded over GPUs run_analysis.sh merge_shards.py ``` `predictions.jsonl` holds one row per sample with the image path, the full prompt, the raw model output, the parsed answer, the reference and correctness (plus per-candidate log-probs under likelihood scoring). ### Running ```bash ./sh/run_all.sh # everything, then the analysis MODEL_ID=Qwen/Qwen2.5-VL-7B-Instruct ./sh/run_sweep.sh vsr ./sh/run_sweep.sh youtube 2 SCORING=likelihood MODEL_ID=Qwen/Qwen3.5-4B-Base ./sh/run_sweep.sh vsr ./sh/run_analysis.sh ``` Everything is resumable: a baseline whose `metrics.json` and `predictions.jsonl` already exist is skipped (`SKIP_COMPLETED=false` to redo). Each baseline runs data-parallel over `GPU_IDS` (default `0 1 2 3`), one process per GPU, shards merged afterwards. Other knobs: `BASELINES`, `BATCH_SIZE`, `MAX_NEW_TOKENS`, `IMAGE_MIN_PIXELS`/`IMAGE_MAX_PIXELS`, `ENABLE_THINKING`, `DEBUG_EVAL`, `PYTHON_BIN`. ### What the analysis reports Per (model, split), over the samples common to all baselines: * **accuracy** per baseline and the spread across the sweep; * **flip rate** — fraction of samples whose *answer* changes relative to baseline 0. This is the direct test of the premise: a visual tower that modelled the scene perfectly would answer identically at every baseline, so the flip rate measures the size of the effect independently of whether the answer was right to begin with; * **stability** — always correct / always wrong / unstable; * **failure overlap** — for every baseline pair, `|F_a ∩ F_b| / |F_a ∪ F_b|` over the failure sets, with raw counts and answer agreement. High overlap means the baselines fail on the *same* samples (a shared difficulty); low overlap means each baseline breaks a different subset, i.e. the failures are driven by the intervention itself. Outputs land in `result/analysis/`: `analysis_summary.json`, one `___failure_overlap.csv` per run, and (with `--dump_per_sample`, which `run_analysis.sh` passes) a per-sample answer/correctness matrix. ### Environment All inference runs in `/home/disheng/miniconda3/envs/Spatial_VLM` (python 3.10, torch 2.8, transformers 5.7). The interpreter is pinned in `sh/env.sh`, which every script sources. It is required, not merely preferred, and there is deliberately **no fallback to `python`**: the base env has transformers 4.51, which has no `qwen3_5` support, so a silent fallback would let the Qwen2.5 sweeps pass and then fail partway through the grid. `sh/env.sh` therefore fails immediately if the interpreter is missing, and `check_env` verifies torch/transformers import and that `qwen3_5` is present before any model loads — once per `run_all.sh`, or once per `run_sweep.sh` when run on its own. It prints, e.g. ``` env ok: python 3.10.19, torch 2.8.0+cu128, transformers 5.7.0, cuda yes ``` Override the interpreter with `PYTHON_BIN=/path/to/python`, or skip the check with `CHECK_ENV=false`.