|
Download readme.md from LLDDSS/Stereo_Depth: direct link, hf CLI and curl.
- Browser
- Download file 9.16 kB
-
https://huggingface.co/LLDDSS/Stereo_Depth/resolve/main/readme.md
- Command line
-
hf download hf://LLDDSS/Stereo_Depth/readme.md
-
curl -L -o readme.md https://huggingface.co/LLDDSS/Stereo_Depth/resolve/main/readme.md
9.16 kB
| this repo used to show the intervention effect of different baseline values. | |
| theorically, if the visual tower can perfectly model the visual world, the small intervention should not change the final VLM output. | |
| this dir is used to check the overlap of failure cases between different baseline values. | |
| Therefore, I will run inference in multiple online model: | |
| 1. Qwen/Qwen2.5-VL-3B-Instruct | |
| 2. Qwen/Qwen2.5-VL-7B-Instruct | |
| 3. Qwen/Qwen3.5-4B | |
| 4. Qwen/Qwen3.5-4B-Base | |
| Datasets: | |
| VSR: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/VSR/V7_testing.json | |
| - baseline: 0 | |
| - baseline: 0.05 | |
| - baseline: 0.1 | |
| - baseline: 0.15 | |
| - baseline: 0.2 | |
| - baseline: 0.25 | |
| Youtube_self_depth_QA | |
| - Level 1: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level1/test.jsonl | |
| - baseline: 0 | |
| - baseline: 0.05 | |
| - baseline: 0.1 | |
| - baseline: 0.15 | |
| - baseline: 0.2 | |
| - baseline: 0.25 | |
| - Level 2: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level2/test.jsonl | |
| - baseline: 0 | |
| - baseline: 0.05 | |
| - baseline: 0.1 | |
| - baseline: 0.15 | |
| - baseline: 0.2 | |
| - baseline: 0.25 | |
| - Level 3: /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/02_data/Youtube_self_depth_QA/level3/test.jsonl | |
| - baseline: 0 | |
| - baseline: 0.05 | |
| - baseline: 0.1 | |
| - baseline: 0.15 | |
| - baseline: 0.2 | |
| - baseline: 0.25 | |
| all of code with respect to load data will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/dataset | |
| all of inference code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/inference | |
| all of inference results will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/results | |
| - it will contain each model's performance in each dataset in each baseline | |
| - each sample's input and output will be saved in the corresponding json file | |
| all of analysis code will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/analysis | |
| all runnable sh scripts will be saved in /HDD2/disheng/Spatial_VLM_Stereo/June/Spatial_VLM_Stereo/01_3_prelimilary_intervention_different_baseline/sh | |
| --- | |
| ## Implementation | |
| At baseline `b` the model sees **one** image: the scene rendered from the | |
| baseline-`b` viewpoint. `b = 0` is the untouched original view, i.e. the | |
| no-intervention control. The question and the reference answer are identical | |
| across baselines, so the per-sample answers line up directly and the failure | |
| sets can be compared. | |
| Baselines are keyed canonically as `0.00 0.05 0.10 0.15 0.20 0.25`. The | |
| on-disk directories are spelled inconsistently (`0.10` vs `0.2`, and the VSR | |
| annotation uses `0.20` for the correspondence but `0.2` for the image); all of | |
| that is normalised in `dataset/baselines.py`, so any spelling can be passed on | |
| the command line. | |
| | split | n | image at baseline b | answer space | | |
| |---|---|---|---| | |
| | VSR (`V7_testing.json`) | 2195 | `00_data/gen_output_VSR/original_images` (b=0) / `.../generate_image/<b>` | True / False | | |
| | Youtube level 1 | 821 | `level1/visualizations/<b>/` | A B C D | | |
| | Youtube level 2 | 764 | `level2/visualizations/<b>/` | A B C D (option letter) | | |
| | Youtube level 3 | 795 | `level3/visualizations/<b>/` | A B C D | | |
| The Youtube visualisations are already rendered per baseline with the | |
| annotation points warped into that view, so the marked points track the scene | |
| across the sweep. | |
| Level 2 is a four-option single choice: the question embeds four candidate | |
| orderings as (A)-(D) and the answer is the option letter. Note that the image | |
| labels its four *points* A-D as well, so a model may answer either with the | |
| option letter ("C") or by writing out that option's ordering ("B > C > A > D"). | |
| Only the letter is the requested format, but the ordering is unambiguous, so it | |
| is mapped back through the sample's own option table rather than scored wrong; | |
| the two are told apart by whether the first standalone letter is followed by a | |
| ">". The loader detects the format from the records themselves, so the older | |
| annotations that asked for the ordering directly (24 candidates, including the | |
| `.ordering_bak` copies) still load and score correctly. | |
| ### Answering modes | |
| `--scoring generate` (default, instruction-tuned models) greedily decodes and | |
| parses the generation. `--scoring likelihood` (base models, which do not follow | |
| the answer format) picks `argmax_c sum_t log P(c_t | prompt, c_<t)` over the | |
| candidate answers, so every sample gets a well-defined answer and nothing | |
| depends on output formatting. When every candidate is a single token -- which | |
| is every split as the data currently stands -- this costs one forward pass per | |
| sample. A multi-token candidate set (the legacy level-2 orderings) falls back to | |
| one forward pass per candidate. | |
| Qwen3.5 is a thinking model and its chat template opens a `<think>` block at | |
| the generation prompt. Thinking is therefore **off** by default | |
| (`--enable_thinking` to turn it on); it is rejected outright under | |
| `--scoring likelihood`, where it would score the candidates as the opening | |
| words of a reasoning trace rather than as the answer. | |
| A generation that never reaches an answer normalises to something outside the | |
| candidate set. Those count as wrong, but are also tallied as `num_unparsed` / | |
| `unparsed_rate` in `metrics.json`, so a high parse-failure rate is visible | |
| instead of being read as low accuracy. If a model's unparsed rate is high, | |
| re-run it with `--scoring likelihood`. | |
| ### Layout | |
| ``` | |
| dataset/ baselines.py canonical baseline naming | |
| datasets.py per-baseline sample loading, prompts, answer parsing | |
| inference/ run_inference.py one model x one split x one baseline | |
| analysis/ analyze_baseline_overlap.py | |
| result/ <model_tag>/<task>/baseline_<b>/{predictions.jsonl,metrics.json,logs/} | |
| sh/ run_all.sh the whole grid (4 models x 4 splits x 6 baselines) | |
| run_sweep.sh one model x one split, all baselines, sharded over GPUs | |
| run_analysis.sh | |
| merge_shards.py | |
| ``` | |
| `predictions.jsonl` holds one row per sample with the image path, the full | |
| prompt, the raw model output, the parsed answer, the reference and correctness | |
| (plus per-candidate log-probs under likelihood scoring). | |
| ### Running | |
| ```bash | |
| ./sh/run_all.sh # everything, then the analysis | |
| MODEL_ID=Qwen/Qwen2.5-VL-7B-Instruct ./sh/run_sweep.sh vsr | |
| ./sh/run_sweep.sh youtube 2 | |
| SCORING=likelihood MODEL_ID=Qwen/Qwen3.5-4B-Base ./sh/run_sweep.sh vsr | |
| ./sh/run_analysis.sh | |
| ``` | |
| Everything is resumable: a baseline whose `metrics.json` and | |
| `predictions.jsonl` already exist is skipped (`SKIP_COMPLETED=false` to redo). | |
| Each baseline runs data-parallel over `GPU_IDS` (default `0 1 2 3`), one | |
| process per GPU, shards merged afterwards. Other knobs: `BASELINES`, | |
| `BATCH_SIZE`, `MAX_NEW_TOKENS`, `IMAGE_MIN_PIXELS`/`IMAGE_MAX_PIXELS`, | |
| `ENABLE_THINKING`, `DEBUG_EVAL`, `PYTHON_BIN`. | |
| ### What the analysis reports | |
| Per (model, split), over the samples common to all baselines: | |
| * **accuracy** per baseline and the spread across the sweep; | |
| * **flip rate** — fraction of samples whose *answer* changes relative to | |
| baseline 0. This is the direct test of the premise: a visual tower that | |
| modelled the scene perfectly would answer identically at every baseline, so | |
| the flip rate measures the size of the effect independently of whether the | |
| answer was right to begin with; | |
| * **stability** — always correct / always wrong / unstable; | |
| * **failure overlap** — for every baseline pair, `|F_a ∩ F_b| / |F_a ∪ F_b|` | |
| over the failure sets, with raw counts and answer agreement. High overlap | |
| means the baselines fail on the *same* samples (a shared difficulty); low | |
| overlap means each baseline breaks a different subset, i.e. the failures are | |
| driven by the intervention itself. | |
| Outputs land in `result/analysis/`: `analysis_summary.json`, one | |
| `<model>__<task>_failure_overlap.csv` per run, and (with `--dump_per_sample`, | |
| which `run_analysis.sh` passes) a per-sample answer/correctness matrix. | |
| ### Environment | |
| All inference runs in `/home/disheng/miniconda3/envs/Spatial_VLM` | |
| (python 3.10, torch 2.8, transformers 5.7). The interpreter is pinned in | |
| `sh/env.sh`, which every script sources. | |
| It is required, not merely preferred, and there is deliberately **no fallback | |
| to `python`**: the base env has transformers 4.51, which has no `qwen3_5` | |
| support, so a silent fallback would let the Qwen2.5 sweeps pass and then fail | |
| partway through the grid. `sh/env.sh` therefore fails immediately if the | |
| interpreter is missing, and `check_env` verifies torch/transformers import and | |
| that `qwen3_5` is present before any model loads — once per `run_all.sh`, or | |
| once per `run_sweep.sh` when run on its own. It prints, e.g. | |
| ``` | |
| env ok: python 3.10.19, torch 2.8.0+cu128, transformers 5.7.0, cuda yes | |
| ``` | |
| Override the interpreter with `PYTHON_BIN=/path/to/python`, or skip the check | |
| with `CHECK_ENV=false`. | |