Ronaldo-GOAT's picture
FORGE3DBench: final eval protocol + eval_final.py, batched Ours inference, held-out view tars, missing-object lists, README
cca6827 verified
|
Raw History Blame Contribute Delete
5.92 kB
# Setup: the drivers are NOT self-contained
Each driver wraps an official repo + its conda env + HF weights. Install each official repo per its README, then point
the drivers at it (edit the path constants at the top of the file, or the env vars below).
| method | official repo @ commit | weights (HF) | driver | notes |
|---|---|---|---|---|
| Pixal3D 1v | TencentARC/Pixal3D @ `cdbb2bb` | `TencentARC/Pixal3D` (single-view `ckpts/*`, `pipeline.json`) | `baselines/pixal3d_ours/batch_pixal3d.py` | copy `mv_common.py`, `mv_proper.py` into the repo root; `PIXAL3D_REPO=<repo>` |
| Pixal3D 4v | TencentARC/Pixal3D @ **`f7cf384`** ("support multi-view image input", 2026-09-01), unmodified | same HF repo, **`ckpts/*_bf16_mv.safetensors` + `pipeline_mv.json`** (+ NAF upsampler, auto-downloaded) | `baselines/pixal3d_ours/batch_pixal3d_mv.py` | uses upstream `inference_mv.py` exactly; `PIXAL3D_UPSTREAM=<f7cf384 checkout>`, `PIXAL3D_OURS=<dir with mv_common.py>`. Same env as Pixal3D 1v (o_voxel, utils3d wheel per README). |
| ReconViaGen | GAP-LAB-CUHK-SZ/ReconViaGen @ `f672092` + `baselines/reconviagen_local.diff` | per its README | `baselines/batch_reconviagen.py` | copy `baselines/reconviagen_ours/infer.py` into the repo root (driver does `from infer import load_pipeline`); edit `REPO`, `ENV`, `HF` |
| Amodal3R | Sm0kyWu/Amodal3R @ `d00e083`, unmodified (ReconViaGen env); driver expects `<AMODAL_ROOT>/repo` | per its README | `baselines/batch_amodal3r.py` | edit `AMODAL_ROOT`, `ENV`, `HF` |
| Hunyuan3D-2mv | Tencent-Hunyuan/Hunyuan3D-2 @ `f8db630`, unmodified; driver expects `<HY_ROOT>/repo`, `<HY_ROOT>/env` | `tencent/Hunyuan3D-2mv` | `baselines/batch_hy3d_2mv.py` | `--seed 42 --simplify-faces 40000`; edit `HY_ROOT`, `HF` |
| Cupid | cupid3d/Cupid @ `10af9b2` | per its README | `baselines/batch_cupid.py` (1v), `baselines/batch_cupid_mv.py` (4v) | see "Cupid 4v" below |
Exact argument lists: `baselines/commands_reference.py` (`cmd_for`). Every driver is skip-existing and takes `--shard i --nshards n`.
Ours: `ours/` needs the training repo (`Ronaldo-GOAT/bert_simpson: migrator/code/mv-sam3d-for-6d-v2-ssflow/`), the SAM3D env, SS-flow 80k + SLAT 32k checkpoints (README.md).
## Our code needed on top of the official repos (all included here)
- ReconViaGen: `baselines/reconviagen_ours/infer.py` (+ `reconviagen_local.diff`, 17 lines in trellis_image_to_3d.py).
- Pixal3D: `baselines/pixal3d_ours/{batch_pixal3d.py, batch_pixal3d_mv.py, mv_common.py, mv_proper.py}`.
- Eval: `eval/eval_final.py` (final evaluator, imports `eval/appeval/*` relative to itself) + `eval/{align_v2.py, cammap_pixal3d.py, cammap_fb150_recenter.py}`; legacy helpers `eval/{faithfulness.py, gt_loader.py, align.py}`.
- Ours: `ours/{dump_seed.py, finish_cache.py, ssflow_coords.py, batch_appforce_sam3d.py, faithfulness.py, gt_loader.py, mkconfig.py, ours_combined_cell.sh, ours32k_queue.sh}`
+ MV-SAM3D (github devinli123/MV-SAM3D @ `abb04b5`; `ours/mv-sam3d_local.diff` touches only run_inference*.py, not used by this path)
+ training repo from the Hub (`migrator/code/mv-sam3d-for-6d-v2-ssflow/`, provides `mvsam3d`, `tools/val_daemon.py`, `configs/train_slatflow_prod.yaml`).
dump_seed/finish_cache import `batch_appforce_sam3d` from `<MVMESH>/metrics` (edit `MVMESH`).
## Batched inference (ours only) = the FINAL inference path
`ours/ours_combined_cell_batched.sh "GPUS" EXP DS NV SSCKPT SLATCKPT SCRATCH TAG` (env `SSB` = stage-1 batch, default 16;
`S2B` = stage-2 batch, default **4**; `OBJFILE` = optional object list; `R`/`REPO`/`PY`/`CPUS` = roots, see the script header).
Pass 2 = `ssflow_coords_batched.py --batch $SSB`, stage 2 = val_daemon with `mkconfig_batched.py` (batch_size=$S2B, single:false).
Keep `S2B <= 4`: with noisy seeds (VGGT / real-photo pointmaps) S2B=8/16 OOMs the batched stage-2 decode on 80 GB GPUs.
Batched vs unbatched (`ours/batch_equivalence.md`): not bit-identical (batch 2 flipped 1 occupancy voxel in 2/4 cases; PSNR diff up
to 0.096 dB, F@0.01 up to 1.8e-3). The FINAL numbers are produced with the batched path (all Ours meshes are regenerated with it,
see README "What needs to be generated"); `ours_combined_cell.sh` (unbatched) is kept as the reference implementation.
Baselines have no batched path (speed-up = several single-object workers per GPU; computation unchanged).
## Pixal3D 4v on FORGE3DBench: re-centred official MV pipeline
`baselines/pixal3d_ours/batch_pixal3d_mv_recenter.py` (+ `recenter.py`) = `batch_pixal3d_mv.py` imported unchanged, only the view
construction replaced: every GT camera is rotated about its own centre so that its optical axis passes through the object centre,
the full frame is warped by the exact homography K R inv(K') (K' centred, cx=cy=S/2, S=1024, object fills 80%), the world is shifted so
the object centre is the origin and gauge-rotated so view 0 is the official canonical front camera. No upstream code is patched.
Run with `--center origin` (object centre = GT-frame origin, which is where the cameras of the benchmark are expressed) and place
the output with `eval/cammap_fb150_recenter.py` (camera-only: X = C'_0 inv(F(d0)) A G + c). Supersedes `batch_pixal3d_mv_cropK.py`.
Toys4K / Omni keep `batch_pixal3d_mv.py` + `eval/cammap_pixal3d.py` (camera-centred renders, no re-centring needed).
## Cupid 4v
`baselines/batch_cupid_mv.py --exp <exp_4v> --out <out> --views 4` (Cupid @ 10af9b2, weights hbb1/Cupid). The paper (Fig. 7, Sec. 6) describes
multi-view as a MultiDiffusion-style test-time fusion of the shared object latent but releases no code: this driver implements that description
(stage 1: shared occupancy latent averaged across per-view flow paths each step, per-view camera latents; stage 2: shared SLAT, per-view
pose-conditioned updates averaged). 1-view mode reproduces the released run() pose exactly. Writes <obj>.pose.json (per-view extrinsics/intrinsics).
~100 s/object, ~15 GB/process.