FORGE3DBench: final eval protocol + eval_final.py, batched Ours inference, held-out view tars, missing-object lists, README
cca6827 verified |
Download forgebench/code/SETUP.md from Ronaldo-GOAT/bert_simpson: direct link, hf CLI and curl.
- Browser
- Download file 5.92 kB
-
https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/forgebench/code/SETUP.md
- Command line
-
hf download hf://Ronaldo-GOAT/bert_simpson/forgebench/code/SETUP.md
-
curl -L -o SETUP.md https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/forgebench/code/SETUP.md
5.92 kB
| # Setup: the drivers are NOT self-contained | |
| Each driver wraps an official repo + its conda env + HF weights. Install each official repo per its README, then point | |
| the drivers at it (edit the path constants at the top of the file, or the env vars below). | |
| | method | official repo @ commit | weights (HF) | driver | notes | | |
| |---|---|---|---|---| | |
| | Pixal3D 1v | TencentARC/Pixal3D @ `cdbb2bb` | `TencentARC/Pixal3D` (single-view `ckpts/*`, `pipeline.json`) | `baselines/pixal3d_ours/batch_pixal3d.py` | copy `mv_common.py`, `mv_proper.py` into the repo root; `PIXAL3D_REPO=<repo>` | | |
| | Pixal3D 4v | TencentARC/Pixal3D @ **`f7cf384`** ("support multi-view image input", 2026-09-01), unmodified | same HF repo, **`ckpts/*_bf16_mv.safetensors` + `pipeline_mv.json`** (+ NAF upsampler, auto-downloaded) | `baselines/pixal3d_ours/batch_pixal3d_mv.py` | uses upstream `inference_mv.py` exactly; `PIXAL3D_UPSTREAM=<f7cf384 checkout>`, `PIXAL3D_OURS=<dir with mv_common.py>`. Same env as Pixal3D 1v (o_voxel, utils3d wheel per README). | | |
| | ReconViaGen | GAP-LAB-CUHK-SZ/ReconViaGen @ `f672092` + `baselines/reconviagen_local.diff` | per its README | `baselines/batch_reconviagen.py` | copy `baselines/reconviagen_ours/infer.py` into the repo root (driver does `from infer import load_pipeline`); edit `REPO`, `ENV`, `HF` | | |
| | Amodal3R | Sm0kyWu/Amodal3R @ `d00e083`, unmodified (ReconViaGen env); driver expects `<AMODAL_ROOT>/repo` | per its README | `baselines/batch_amodal3r.py` | edit `AMODAL_ROOT`, `ENV`, `HF` | | |
| | Hunyuan3D-2mv | Tencent-Hunyuan/Hunyuan3D-2 @ `f8db630`, unmodified; driver expects `<HY_ROOT>/repo`, `<HY_ROOT>/env` | `tencent/Hunyuan3D-2mv` | `baselines/batch_hy3d_2mv.py` | `--seed 42 --simplify-faces 40000`; edit `HY_ROOT`, `HF` | | |
| | Cupid | cupid3d/Cupid @ `10af9b2` | per its README | `baselines/batch_cupid.py` (1v), `baselines/batch_cupid_mv.py` (4v) | see "Cupid 4v" below | | |
| Exact argument lists: `baselines/commands_reference.py` (`cmd_for`). Every driver is skip-existing and takes `--shard i --nshards n`. | |
| Ours: `ours/` needs the training repo (`Ronaldo-GOAT/bert_simpson: migrator/code/mv-sam3d-for-6d-v2-ssflow/`), the SAM3D env, SS-flow 80k + SLAT 32k checkpoints (README.md). | |
| ## Our code needed on top of the official repos (all included here) | |
| - ReconViaGen: `baselines/reconviagen_ours/infer.py` (+ `reconviagen_local.diff`, 17 lines in trellis_image_to_3d.py). | |
| - Pixal3D: `baselines/pixal3d_ours/{batch_pixal3d.py, batch_pixal3d_mv.py, mv_common.py, mv_proper.py}`. | |
| - Eval: `eval/eval_final.py` (final evaluator, imports `eval/appeval/*` relative to itself) + `eval/{align_v2.py, cammap_pixal3d.py, cammap_fb150_recenter.py}`; legacy helpers `eval/{faithfulness.py, gt_loader.py, align.py}`. | |
| - Ours: `ours/{dump_seed.py, finish_cache.py, ssflow_coords.py, batch_appforce_sam3d.py, faithfulness.py, gt_loader.py, mkconfig.py, ours_combined_cell.sh, ours32k_queue.sh}` | |
| + MV-SAM3D (github devinli123/MV-SAM3D @ `abb04b5`; `ours/mv-sam3d_local.diff` touches only run_inference*.py, not used by this path) | |
| + training repo from the Hub (`migrator/code/mv-sam3d-for-6d-v2-ssflow/`, provides `mvsam3d`, `tools/val_daemon.py`, `configs/train_slatflow_prod.yaml`). | |
| dump_seed/finish_cache import `batch_appforce_sam3d` from `<MVMESH>/metrics` (edit `MVMESH`). | |
| ## Batched inference (ours only) = the FINAL inference path | |
| `ours/ours_combined_cell_batched.sh "GPUS" EXP DS NV SSCKPT SLATCKPT SCRATCH TAG` (env `SSB` = stage-1 batch, default 16; | |
| `S2B` = stage-2 batch, default **4**; `OBJFILE` = optional object list; `R`/`REPO`/`PY`/`CPUS` = roots, see the script header). | |
| Pass 2 = `ssflow_coords_batched.py --batch $SSB`, stage 2 = val_daemon with `mkconfig_batched.py` (batch_size=$S2B, single:false). | |
| Keep `S2B <= 4`: with noisy seeds (VGGT / real-photo pointmaps) S2B=8/16 OOMs the batched stage-2 decode on 80 GB GPUs. | |
| Batched vs unbatched (`ours/batch_equivalence.md`): not bit-identical (batch 2 flipped 1 occupancy voxel in 2/4 cases; PSNR diff up | |
| to 0.096 dB, F@0.01 up to 1.8e-3). The FINAL numbers are produced with the batched path (all Ours meshes are regenerated with it, | |
| see README "What needs to be generated"); `ours_combined_cell.sh` (unbatched) is kept as the reference implementation. | |
| Baselines have no batched path (speed-up = several single-object workers per GPU; computation unchanged). | |
| ## Pixal3D 4v on FORGE3DBench: re-centred official MV pipeline | |
| `baselines/pixal3d_ours/batch_pixal3d_mv_recenter.py` (+ `recenter.py`) = `batch_pixal3d_mv.py` imported unchanged, only the view | |
| construction replaced: every GT camera is rotated about its own centre so that its optical axis passes through the object centre, | |
| the full frame is warped by the exact homography K R inv(K') (K' centred, cx=cy=S/2, S=1024, object fills 80%), the world is shifted so | |
| the object centre is the origin and gauge-rotated so view 0 is the official canonical front camera. No upstream code is patched. | |
| Run with `--center origin` (object centre = GT-frame origin, which is where the cameras of the benchmark are expressed) and place | |
| the output with `eval/cammap_fb150_recenter.py` (camera-only: X = C'_0 inv(F(d0)) A G + c). Supersedes `batch_pixal3d_mv_cropK.py`. | |
| Toys4K / Omni keep `batch_pixal3d_mv.py` + `eval/cammap_pixal3d.py` (camera-centred renders, no re-centring needed). | |
| ## Cupid 4v | |
| `baselines/batch_cupid_mv.py --exp <exp_4v> --out <out> --views 4` (Cupid @ 10af9b2, weights hbb1/Cupid). The paper (Fig. 7, Sec. 6) describes | |
| multi-view as a MultiDiffusion-style test-time fusion of the shared object latent but releases no code: this driver implements that description | |
| (stage 1: shared occupancy latent averaged across per-view flow paths each step, per-view camera latents; stage 2: shared SLAT, per-view | |
| pose-conditioned updates averaged). 1-view mode reproduces the released run() pose exactly. Writes <obj>.pose.json (per-view extrinsics/intrinsics). | |
| ~100 s/object, ~15 GB/process. | |