Ronaldo-GOAT's picture
FORGE3DBench: final eval protocol + eval_final.py, batched Ours inference, held-out view tars, missing-object lists, README
cca6827 verified
|
Raw History Blame Contribute Delete
5.92 kB

Setup: the drivers are NOT self-contained

Each driver wraps an official repo + its conda env + HF weights. Install each official repo per its README, then point the drivers at it (edit the path constants at the top of the file, or the env vars below).

method official repo @ commit weights (HF) driver notes
Pixal3D 1v TencentARC/Pixal3D @ cdbb2bb TencentARC/Pixal3D (single-view ckpts/*, pipeline.json) baselines/pixal3d_ours/batch_pixal3d.py copy mv_common.py, mv_proper.py into the repo root; PIXAL3D_REPO=<repo>
Pixal3D 4v TencentARC/Pixal3D @ f7cf384 ("support multi-view image input", 2026-09-01), unmodified same HF repo, ckpts/*_bf16_mv.safetensors + pipeline_mv.json (+ NAF upsampler, auto-downloaded) baselines/pixal3d_ours/batch_pixal3d_mv.py uses upstream inference_mv.py exactly; PIXAL3D_UPSTREAM=<f7cf384 checkout>, PIXAL3D_OURS=<dir with mv_common.py>. Same env as Pixal3D 1v (o_voxel, utils3d wheel per README).
ReconViaGen GAP-LAB-CUHK-SZ/ReconViaGen @ f672092 + baselines/reconviagen_local.diff per its README baselines/batch_reconviagen.py copy baselines/reconviagen_ours/infer.py into the repo root (driver does from infer import load_pipeline); edit REPO, ENV, HF
Amodal3R Sm0kyWu/Amodal3R @ d00e083, unmodified (ReconViaGen env); driver expects <AMODAL_ROOT>/repo per its README baselines/batch_amodal3r.py edit AMODAL_ROOT, ENV, HF
Hunyuan3D-2mv Tencent-Hunyuan/Hunyuan3D-2 @ f8db630, unmodified; driver expects <HY_ROOT>/repo, <HY_ROOT>/env tencent/Hunyuan3D-2mv baselines/batch_hy3d_2mv.py --seed 42 --simplify-faces 40000; edit HY_ROOT, HF
Cupid cupid3d/Cupid @ 10af9b2 per its README baselines/batch_cupid.py (1v), baselines/batch_cupid_mv.py (4v) see "Cupid 4v" below

Exact argument lists: baselines/commands_reference.py (cmd_for). Every driver is skip-existing and takes --shard i --nshards n. Ours: ours/ needs the training repo (Ronaldo-GOAT/bert_simpson: migrator/code/mv-sam3d-for-6d-v2-ssflow/), the SAM3D env, SS-flow 80k + SLAT 32k checkpoints (README.md).

Our code needed on top of the official repos (all included here)

  • ReconViaGen: baselines/reconviagen_ours/infer.py (+ reconviagen_local.diff, 17 lines in trellis_image_to_3d.py).
  • Pixal3D: baselines/pixal3d_ours/{batch_pixal3d.py, batch_pixal3d_mv.py, mv_common.py, mv_proper.py}.
  • Eval: eval/eval_final.py (final evaluator, imports eval/appeval/* relative to itself) + eval/{align_v2.py, cammap_pixal3d.py, cammap_fb150_recenter.py}; legacy helpers eval/{faithfulness.py, gt_loader.py, align.py}.
  • Ours: ours/{dump_seed.py, finish_cache.py, ssflow_coords.py, batch_appforce_sam3d.py, faithfulness.py, gt_loader.py, mkconfig.py, ours_combined_cell.sh, ours32k_queue.sh}
    • MV-SAM3D (github devinli123/MV-SAM3D @ abb04b5; ours/mv-sam3d_local.diff touches only run_inference*.py, not used by this path)
    • training repo from the Hub (migrator/code/mv-sam3d-for-6d-v2-ssflow/, provides mvsam3d, tools/val_daemon.py, configs/train_slatflow_prod.yaml). dump_seed/finish_cache import batch_appforce_sam3d from <MVMESH>/metrics (edit MVMESH).

Batched inference (ours only) = the FINAL inference path

ours/ours_combined_cell_batched.sh "GPUS" EXP DS NV SSCKPT SLATCKPT SCRATCH TAG (env SSB = stage-1 batch, default 16; S2B = stage-2 batch, default 4; OBJFILE = optional object list; R/REPO/PY/CPUS = roots, see the script header). Pass 2 = ssflow_coords_batched.py --batch $SSB, stage 2 = val_daemon with mkconfig_batched.py (batch_size=$S2B, single:false). Keep S2B <= 4: with noisy seeds (VGGT / real-photo pointmaps) S2B=8/16 OOMs the batched stage-2 decode on 80 GB GPUs. Batched vs unbatched (ours/batch_equivalence.md): not bit-identical (batch 2 flipped 1 occupancy voxel in 2/4 cases; PSNR diff up to 0.096 dB, F@0.01 up to 1.8e-3). The FINAL numbers are produced with the batched path (all Ours meshes are regenerated with it, see README "What needs to be generated"); ours_combined_cell.sh (unbatched) is kept as the reference implementation. Baselines have no batched path (speed-up = several single-object workers per GPU; computation unchanged).

Pixal3D 4v on FORGE3DBench: re-centred official MV pipeline

baselines/pixal3d_ours/batch_pixal3d_mv_recenter.py (+ recenter.py) = batch_pixal3d_mv.py imported unchanged, only the view construction replaced: every GT camera is rotated about its own centre so that its optical axis passes through the object centre, the full frame is warped by the exact homography K R inv(K') (K' centred, cx=cy=S/2, S=1024, object fills 80%), the world is shifted so the object centre is the origin and gauge-rotated so view 0 is the official canonical front camera. No upstream code is patched. Run with --center origin (object centre = GT-frame origin, which is where the cameras of the benchmark are expressed) and place the output with eval/cammap_fb150_recenter.py (camera-only: X = C'_0 inv(F(d0)) A G + c). Supersedes batch_pixal3d_mv_cropK.py. Toys4K / Omni keep batch_pixal3d_mv.py + eval/cammap_pixal3d.py (camera-centred renders, no re-centring needed).

Cupid 4v

baselines/batch_cupid_mv.py --exp <exp_4v> --out <out> --views 4 (Cupid @ 10af9b2, weights hbb1/Cupid). The paper (Fig. 7, Sec. 6) describes multi-view as a MultiDiffusion-style test-time fusion of the shared object latent but releases no code: this driver implements that description (stage 1: shared occupancy latent averaged across per-view flow paths each step, per-view camera latents; stage 2: shared SLAT, per-view pose-conditioned updates averaged). 1-view mode reproduces the released run() pose exactly. Writes .pose.json (per-view extrinsics/intrinsics). ~100 s/object, ~15 GB/process.