Stereo_Depth / readme.md
LLDDSS's picture
Upload folder using huggingface_hub
18104f9 verified
|
Raw History Blame Contribute Delete
9.16 kB

this repo used to show the intervention effect of different baseline values.

theorically, if the visual tower can perfectly model the visual world, the small intervention should not change the final VLM output.

this dir is used to check the overlap of failure cases between different baseline values.

Therefore, I will run inference in multiple online model:

  1. Qwen/Qwen2.5-VL-3B-Instruct

  2. Qwen/Qwen2.5-VL-7B-Instruct

  3. Qwen/Qwen3.5-4B

  4. Qwen/Qwen3.5-4B-Base

Datasets: VSR: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/VSR/V7_testing.json

  • baseline: 0
  • baseline: 0.05
  • baseline: 0.1
  • baseline: 0.15
  • baseline: 0.2
  • baseline: 0.25

Youtube_self_depth_QA

  • Level 1: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level1/test.jsonl

    • baseline: 0
    • baseline: 0.05
    • baseline: 0.1
    • baseline: 0.15
    • baseline: 0.2
    • baseline: 0.25
  • Level 2: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level2/test.jsonl

    • baseline: 0
    • baseline: 0.05
    • baseline: 0.1
    • baseline: 0.15
    • baseline: 0.2
    • baseline: 0.25
  • Level 3: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level3/test.jsonl

    • baseline: 0
    • baseline: 0.05
    • baseline: 0.1
    • baseline: 0.15
    • baseline: 0.2
    • baseline: 0.25

all of code with respect to load data will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/dataset

all of inference code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/inference

all of inference results will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/results - it will contain each model's performance in each dataset in each baseline - each sample's input and output will be saved in the corresponding json file

all of analysis code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/analysis

all runnable sh scripts will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/sh


Implementation

At baseline b the model sees one image: the scene rendered from the baseline-b viewpoint. b = 0 is the untouched original view, i.e. the no-intervention control. The question and the reference answer are identical across baselines, so the per-sample answers line up directly and the failure sets can be compared.

Baselines are keyed canonically as 0.00 0.05 0.10 0.15 0.20 0.25. The on-disk directories are spelled inconsistently (0.10 vs 0.2, and the VSR annotation uses 0.20 for the correspondence but 0.2 for the image); all of that is normalised in dataset/baselines.py, so any spelling can be passed on the command line.

split n image at baseline b answer space
VSR (V7_testing.json) 2195 00_data/gen_output_VSR/original_images (b=0) / .../generate_image/<b> True / False
Youtube level 1 821 level1/visualizations/<b>/ A B C D
Youtube level 2 764 level2/visualizations/<b>/ A B C D (option letter)
Youtube level 3 795 level3/visualizations/<b>/ A B C D

The Youtube visualisations are already rendered per baseline with the annotation points warped into that view, so the marked points track the scene across the sweep.

Level 2 is a four-option single choice: the question embeds four candidate orderings as (A)-(D) and the answer is the option letter. Note that the image labels its four points A-D as well, so a model may answer either with the option letter ("C") or by writing out that option's ordering ("B > C > A > D"). Only the letter is the requested format, but the ordering is unambiguous, so it is mapped back through the sample's own option table rather than scored wrong; the two are told apart by whether the first standalone letter is followed by a ">". The loader detects the format from the records themselves, so the older annotations that asked for the ordering directly (24 candidates, including the .ordering_bak copies) still load and score correctly.

Answering modes

--scoring generate (default, instruction-tuned models) greedily decodes and parses the generation. --scoring likelihood (base models, which do not follow the answer format) picks argmax_c sum_t log P(c_t | prompt, c_<t) over the candidate answers, so every sample gets a well-defined answer and nothing depends on output formatting. When every candidate is a single token -- which is every split as the data currently stands -- this costs one forward pass per sample. A multi-token candidate set (the legacy level-2 orderings) falls back to one forward pass per candidate.

Qwen3.5 is a thinking model and its chat template opens a <think> block at the generation prompt. Thinking is therefore off by default (--enable_thinking to turn it on); it is rejected outright under --scoring likelihood, where it would score the candidates as the opening words of a reasoning trace rather than as the answer.

A generation that never reaches an answer normalises to something outside the candidate set. Those count as wrong, but are also tallied as num_unparsed / unparsed_rate in metrics.json, so a high parse-failure rate is visible instead of being read as low accuracy. If a model's unparsed rate is high, re-run it with --scoring likelihood.

Layout

dataset/    baselines.py  canonical baseline naming
            datasets.py   per-baseline sample loading, prompts, answer parsing
inference/  run_inference.py   one model x one split x one baseline
analysis/   analyze_baseline_overlap.py
result/     <model_tag>/<task>/baseline_<b>/{predictions.jsonl,metrics.json,logs/}
sh/         run_all.sh     the whole grid (4 models x 4 splits x 6 baselines)
            run_sweep.sh   one model x one split, all baselines, sharded over GPUs
            run_analysis.sh
            merge_shards.py

predictions.jsonl holds one row per sample with the image path, the full prompt, the raw model output, the parsed answer, the reference and correctness (plus per-candidate log-probs under likelihood scoring).

Running

./sh/run_all.sh                                     # everything, then the analysis

MODEL_ID=Qwen/Qwen2.5-VL-7B-Instruct ./sh/run_sweep.sh vsr
./sh/run_sweep.sh youtube 2
SCORING=likelihood MODEL_ID=Qwen/Qwen3.5-4B-Base ./sh/run_sweep.sh vsr

./sh/run_analysis.sh

Everything is resumable: a baseline whose metrics.json and predictions.jsonl already exist is skipped (SKIP_COMPLETED=false to redo). Each baseline runs data-parallel over GPU_IDS (default 0 1 2 3), one process per GPU, shards merged afterwards. Other knobs: BASELINES, BATCH_SIZE, MAX_NEW_TOKENS, IMAGE_MIN_PIXELS/IMAGE_MAX_PIXELS, ENABLE_THINKING, DEBUG_EVAL, PYTHON_BIN.

What the analysis reports

Per (model, split), over the samples common to all baselines:

  • accuracy per baseline and the spread across the sweep;
  • flip rate — fraction of samples whose answer changes relative to baseline 0. This is the direct test of the premise: a visual tower that modelled the scene perfectly would answer identically at every baseline, so the flip rate measures the size of the effect independently of whether the answer was right to begin with;
  • stability — always correct / always wrong / unstable;
  • failure overlap — for every baseline pair, |F_a ∩ F_b| / |F_a ∪ F_b| over the failure sets, with raw counts and answer agreement. High overlap means the baselines fail on the same samples (a shared difficulty); low overlap means each baseline breaks a different subset, i.e. the failures are driven by the intervention itself.

Outputs land in result/analysis/: analysis_summary.json, one <model>__<task>_failure_overlap.csv per run, and (with --dump_per_sample, which run_analysis.sh passes) a per-sample answer/correctness matrix.

Environment

All inference runs in /home/disheng/miniconda3/envs/Spatial_VLM (python 3.10, torch 2.8, transformers 5.7). The interpreter is pinned in sh/env.sh, which every script sources.

It is required, not merely preferred, and there is deliberately no fallback to python: the base env has transformers 4.51, which has no qwen3_5 support, so a silent fallback would let the Qwen2.5 sweeps pass and then fail partway through the grid. sh/env.sh therefore fails immediately if the interpreter is missing, and check_env verifies torch/transformers import and that qwen3_5 is present before any model loads — once per run_all.sh, or once per run_sweep.sh when run on its own. It prints, e.g.

env ok: python 3.10.19, torch 2.8.0+cu128, transformers 5.7.0, cuda yes

Override the interpreter with PYTHON_BIN=/path/to/python, or skip the check with CHECK_ENV=false.