Download readme.md from LLDDSS/Stereo_Depth: direct link, hf CLI and curl.
- Browser
- Download file 9.16 kB
-
https://huggingface.co/LLDDSS/Stereo_Depth/resolve/main/readme.md
- Command line
-
hf download hf://LLDDSS/Stereo_Depth/readme.md
-
curl -L -o readme.md https://huggingface.co/LLDDSS/Stereo_Depth/resolve/main/readme.md
this repo used to show the intervention effect of different baseline values.
theorically, if the visual tower can perfectly model the visual world, the small intervention should not change the final VLM output.
this dir is used to check the overlap of failure cases between different baseline values.
Therefore, I will run inference in multiple online model:
Qwen/Qwen2.5-VL-3B-Instruct
Qwen/Qwen2.5-VL-7B-Instruct
Qwen/Qwen3.5-4B
Qwen/Qwen3.5-4B-Base
Datasets: VSR: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/VSR/V7_testing.json
- baseline: 0
- baseline: 0.05
- baseline: 0.1
- baseline: 0.15
- baseline: 0.2
- baseline: 0.25
Youtube_self_depth_QA
Level 1: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level1/test.jsonl
- baseline: 0
- baseline: 0.05
- baseline: 0.1
- baseline: 0.15
- baseline: 0.2
- baseline: 0.25
Level 2: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level2/test.jsonl
- baseline: 0
- baseline: 0.05
- baseline: 0.1
- baseline: 0.15
- baseline: 0.2
- baseline: 0.25
Level 3: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level3/test.jsonl
- baseline: 0
- baseline: 0.05
- baseline: 0.1
- baseline: 0.15
- baseline: 0.2
- baseline: 0.25
all of code with respect to load data will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/dataset
all of inference code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/inference
all of inference results will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/results - it will contain each model's performance in each dataset in each baseline - each sample's input and output will be saved in the corresponding json file
all of analysis code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/analysis
all runnable sh scripts will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/sh
Implementation
At baseline b the model sees one image: the scene rendered from the
baseline-b viewpoint. b = 0 is the untouched original view, i.e. the
no-intervention control. The question and the reference answer are identical
across baselines, so the per-sample answers line up directly and the failure
sets can be compared.
Baselines are keyed canonically as 0.00 0.05 0.10 0.15 0.20 0.25. The
on-disk directories are spelled inconsistently (0.10 vs 0.2, and the VSR
annotation uses 0.20 for the correspondence but 0.2 for the image); all of
that is normalised in dataset/baselines.py, so any spelling can be passed on
the command line.
| split | n | image at baseline b | answer space |
|---|---|---|---|
VSR (V7_testing.json) |
2195 | 00_data/gen_output_VSR/original_images (b=0) / .../generate_image/<b> |
True / False |
| Youtube level 1 | 821 | level1/visualizations/<b>/ |
A B C D |
| Youtube level 2 | 764 | level2/visualizations/<b>/ |
A B C D (option letter) |
| Youtube level 3 | 795 | level3/visualizations/<b>/ |
A B C D |
The Youtube visualisations are already rendered per baseline with the annotation points warped into that view, so the marked points track the scene across the sweep.
Level 2 is a four-option single choice: the question embeds four candidate
orderings as (A)-(D) and the answer is the option letter. Note that the image
labels its four points A-D as well, so a model may answer either with the
option letter ("C") or by writing out that option's ordering ("B > C > A > D").
Only the letter is the requested format, but the ordering is unambiguous, so it
is mapped back through the sample's own option table rather than scored wrong;
the two are told apart by whether the first standalone letter is followed by a
">". The loader detects the format from the records themselves, so the older
annotations that asked for the ordering directly (24 candidates, including the
.ordering_bak copies) still load and score correctly.
Answering modes
--scoring generate (default, instruction-tuned models) greedily decodes and
parses the generation. --scoring likelihood (base models, which do not follow
the answer format) picks argmax_c sum_t log P(c_t | prompt, c_<t) over the
candidate answers, so every sample gets a well-defined answer and nothing
depends on output formatting. When every candidate is a single token -- which
is every split as the data currently stands -- this costs one forward pass per
sample. A multi-token candidate set (the legacy level-2 orderings) falls back to
one forward pass per candidate.
Qwen3.5 is a thinking model and its chat template opens a <think> block at
the generation prompt. Thinking is therefore off by default
(--enable_thinking to turn it on); it is rejected outright under
--scoring likelihood, where it would score the candidates as the opening
words of a reasoning trace rather than as the answer.
A generation that never reaches an answer normalises to something outside the
candidate set. Those count as wrong, but are also tallied as num_unparsed /
unparsed_rate in metrics.json, so a high parse-failure rate is visible
instead of being read as low accuracy. If a model's unparsed rate is high,
re-run it with --scoring likelihood.
Layout
dataset/ baselines.py canonical baseline naming
datasets.py per-baseline sample loading, prompts, answer parsing
inference/ run_inference.py one model x one split x one baseline
analysis/ analyze_baseline_overlap.py
result/ <model_tag>/<task>/baseline_<b>/{predictions.jsonl,metrics.json,logs/}
sh/ run_all.sh the whole grid (4 models x 4 splits x 6 baselines)
run_sweep.sh one model x one split, all baselines, sharded over GPUs
run_analysis.sh
merge_shards.py
predictions.jsonl holds one row per sample with the image path, the full
prompt, the raw model output, the parsed answer, the reference and correctness
(plus per-candidate log-probs under likelihood scoring).
Running
./sh/run_all.sh # everything, then the analysis
MODEL_ID=Qwen/Qwen2.5-VL-7B-Instruct ./sh/run_sweep.sh vsr
./sh/run_sweep.sh youtube 2
SCORING=likelihood MODEL_ID=Qwen/Qwen3.5-4B-Base ./sh/run_sweep.sh vsr
./sh/run_analysis.sh
Everything is resumable: a baseline whose metrics.json and
predictions.jsonl already exist is skipped (SKIP_COMPLETED=false to redo).
Each baseline runs data-parallel over GPU_IDS (default 0 1 2 3), one
process per GPU, shards merged afterwards. Other knobs: BASELINES,
BATCH_SIZE, MAX_NEW_TOKENS, IMAGE_MIN_PIXELS/IMAGE_MAX_PIXELS,
ENABLE_THINKING, DEBUG_EVAL, PYTHON_BIN.
What the analysis reports
Per (model, split), over the samples common to all baselines:
- accuracy per baseline and the spread across the sweep;
- flip rate — fraction of samples whose answer changes relative to baseline 0. This is the direct test of the premise: a visual tower that modelled the scene perfectly would answer identically at every baseline, so the flip rate measures the size of the effect independently of whether the answer was right to begin with;
- stability — always correct / always wrong / unstable;
- failure overlap — for every baseline pair,
|F_a ∩ F_b| / |F_a ∪ F_b|over the failure sets, with raw counts and answer agreement. High overlap means the baselines fail on the same samples (a shared difficulty); low overlap means each baseline breaks a different subset, i.e. the failures are driven by the intervention itself.
Outputs land in result/analysis/: analysis_summary.json, one
<model>__<task>_failure_overlap.csv per run, and (with --dump_per_sample,
which run_analysis.sh passes) a per-sample answer/correctness matrix.
Environment
All inference runs in /home/disheng/miniconda3/envs/Spatial_VLM
(python 3.10, torch 2.8, transformers 5.7). The interpreter is pinned in
sh/env.sh, which every script sources.
It is required, not merely preferred, and there is deliberately no fallback
to python: the base env has transformers 4.51, which has no qwen3_5
support, so a silent fallback would let the Qwen2.5 sweeps pass and then fail
partway through the grid. sh/env.sh therefore fails immediately if the
interpreter is missing, and check_env verifies torch/transformers import and
that qwen3_5 is present before any model loads — once per run_all.sh, or
once per run_sweep.sh when run on its own. It prints, e.g.
env ok: python 3.10.19, torch 2.8.0+cu128, transformers 5.7.0, cuda yes
Override the interpreter with PYTHON_BIN=/path/to/python, or skip the check
with CHECK_ENV=false.