# Setup: the drivers are NOT self-contained Each driver wraps an official repo + its conda env + HF weights. Install each official repo per its README, then point the drivers at it (edit the path constants at the top of the file, or the env vars below). | method | official repo @ commit | weights (HF) | driver | notes | |---|---|---|---|---| | Pixal3D 1v | TencentARC/Pixal3D @ `cdbb2bb` | `TencentARC/Pixal3D` (single-view `ckpts/*`, `pipeline.json`) | `baselines/pixal3d_ours/batch_pixal3d.py` | copy `mv_common.py`, `mv_proper.py` into the repo root; `PIXAL3D_REPO=` | | Pixal3D 4v | TencentARC/Pixal3D @ **`f7cf384`** ("support multi-view image input", 2026-09-01), unmodified | same HF repo, **`ckpts/*_bf16_mv.safetensors` + `pipeline_mv.json`** (+ NAF upsampler, auto-downloaded) | `baselines/pixal3d_ours/batch_pixal3d_mv.py` | uses upstream `inference_mv.py` exactly; `PIXAL3D_UPSTREAM=`, `PIXAL3D_OURS=`. Same env as Pixal3D 1v (o_voxel, utils3d wheel per README). | | ReconViaGen | GAP-LAB-CUHK-SZ/ReconViaGen @ `f672092` + `baselines/reconviagen_local.diff` | per its README | `baselines/batch_reconviagen.py` | copy `baselines/reconviagen_ours/infer.py` into the repo root (driver does `from infer import load_pipeline`); edit `REPO`, `ENV`, `HF` | | Amodal3R | Sm0kyWu/Amodal3R @ `d00e083`, unmodified (ReconViaGen env); driver expects `/repo` | per its README | `baselines/batch_amodal3r.py` | edit `AMODAL_ROOT`, `ENV`, `HF` | | Hunyuan3D-2mv | Tencent-Hunyuan/Hunyuan3D-2 @ `f8db630`, unmodified; driver expects `/repo`, `/env` | `tencent/Hunyuan3D-2mv` | `baselines/batch_hy3d_2mv.py` | `--seed 42 --simplify-faces 40000`; edit `HY_ROOT`, `HF` | | Cupid | cupid3d/Cupid @ `10af9b2` | per its README | `baselines/batch_cupid.py` (1v), `baselines/batch_cupid_mv.py` (4v) | see "Cupid 4v" below | Exact argument lists: `baselines/commands_reference.py` (`cmd_for`). Every driver is skip-existing and takes `--shard i --nshards n`. Ours: `ours/` needs the training repo (`Ronaldo-GOAT/bert_simpson: migrator/code/mv-sam3d-for-6d-v2-ssflow/`), the SAM3D env, SS-flow 80k + SLAT 32k checkpoints (README.md). ## Our code needed on top of the official repos (all included here) - ReconViaGen: `baselines/reconviagen_ours/infer.py` (+ `reconviagen_local.diff`, 17 lines in trellis_image_to_3d.py). - Pixal3D: `baselines/pixal3d_ours/{batch_pixal3d.py, batch_pixal3d_mv.py, mv_common.py, mv_proper.py}`. - Eval: `eval/eval_final.py` (final evaluator, imports `eval/appeval/*` relative to itself) + `eval/{align_v2.py, cammap_pixal3d.py, cammap_fb150_recenter.py}`; legacy helpers `eval/{faithfulness.py, gt_loader.py, align.py}`. - Ours: `ours/{dump_seed.py, finish_cache.py, ssflow_coords.py, batch_appforce_sam3d.py, faithfulness.py, gt_loader.py, mkconfig.py, ours_combined_cell.sh, ours32k_queue.sh}` + MV-SAM3D (github devinli123/MV-SAM3D @ `abb04b5`; `ours/mv-sam3d_local.diff` touches only run_inference*.py, not used by this path) + training repo from the Hub (`migrator/code/mv-sam3d-for-6d-v2-ssflow/`, provides `mvsam3d`, `tools/val_daemon.py`, `configs/train_slatflow_prod.yaml`). dump_seed/finish_cache import `batch_appforce_sam3d` from `/metrics` (edit `MVMESH`). ## Batched inference (ours only) = the FINAL inference path `ours/ours_combined_cell_batched.sh "GPUS" EXP DS NV SSCKPT SLATCKPT SCRATCH TAG` (env `SSB` = stage-1 batch, default 16; `S2B` = stage-2 batch, default **4**; `OBJFILE` = optional object list; `R`/`REPO`/`PY`/`CPUS` = roots, see the script header). Pass 2 = `ssflow_coords_batched.py --batch $SSB`, stage 2 = val_daemon with `mkconfig_batched.py` (batch_size=$S2B, single:false). Keep `S2B <= 4`: with noisy seeds (VGGT / real-photo pointmaps) S2B=8/16 OOMs the batched stage-2 decode on 80 GB GPUs. Batched vs unbatched (`ours/batch_equivalence.md`): not bit-identical (batch 2 flipped 1 occupancy voxel in 2/4 cases; PSNR diff up to 0.096 dB, F@0.01 up to 1.8e-3). The FINAL numbers are produced with the batched path (all Ours meshes are regenerated with it, see README "What needs to be generated"); `ours_combined_cell.sh` (unbatched) is kept as the reference implementation. Baselines have no batched path (speed-up = several single-object workers per GPU; computation unchanged). ## Pixal3D 4v on FORGE3DBench: re-centred official MV pipeline `baselines/pixal3d_ours/batch_pixal3d_mv_recenter.py` (+ `recenter.py`) = `batch_pixal3d_mv.py` imported unchanged, only the view construction replaced: every GT camera is rotated about its own centre so that its optical axis passes through the object centre, the full frame is warped by the exact homography K R inv(K') (K' centred, cx=cy=S/2, S=1024, object fills 80%), the world is shifted so the object centre is the origin and gauge-rotated so view 0 is the official canonical front camera. No upstream code is patched. Run with `--center origin` (object centre = GT-frame origin, which is where the cameras of the benchmark are expressed) and place the output with `eval/cammap_fb150_recenter.py` (camera-only: X = C'_0 inv(F(d0)) A G + c). Supersedes `batch_pixal3d_mv_cropK.py`. Toys4K / Omni keep `batch_pixal3d_mv.py` + `eval/cammap_pixal3d.py` (camera-centred renders, no re-centring needed). ## Cupid 4v `baselines/batch_cupid_mv.py --exp --out --views 4` (Cupid @ 10af9b2, weights hbb1/Cupid). The paper (Fig. 7, Sec. 6) describes multi-view as a MultiDiffusion-style test-time fusion of the shared object latent but releases no code: this driver implements that description (stage 1: shared occupancy latent averaged across per-view flow paths each step, per-view camera latents; stage 2: shared SLAT, per-view pose-conditioned updates averaged). 1-view mode reproduces the released run() pose exactly. Writes .pose.json (per-view extrinsics/intrinsics). ~100 s/object, ~15 GB/process.