# DOWNLOAD_GUIDE — per-dataset acquisition + conversion How to obtain each dataset we integrated and convert it into the unified contract (`DATA_FORMAT.md`). For each dataset: **host/URL**, **access** (open vs gated), **approx size**, **how to get it**, and the **coordinate conversion** to our contract (condensed from the verified `code/mv-sam3d-for-6d-v2/docs/dataset_convention_spec.md` — that file is the authority if anything here is ambiguous). Put everything under one root and point the config at it: ```bash export MIGRATOR_DATA_ROOT=/path/to/datasets # each dataset in its own subdir ``` After downloading a dataset, build its canonical SS latents once (see `DATA_FORMAT.md` §4) before training on it. Legend for conversion status (from the spec): ✅ = empirically reprojection-IoU verified; 🔵 = code/cross-check verified. Notation: `R_oc,t_oc` = object(CAD)→cam OpenCV meters (`x_cam=R_oc·p_cad+t_oc`); `p_cad` = mesh verts (m). Adapters return raw OpenCV meters — never pre-apply the `D=diag(-1,-1,1)` flip or the canonical `M` (the base loader does that). --- # TRAINING DATASETS ## DexYCB (reference; 20 YCB objects) — subdir `dexycb/` - **Host**: https://dex-ycb.github.io/ (NVIDIA). Open, registration-free download links on the project page ("Download"). - **Size**: ~1.4 TB full; the s0/s1 subset + models is enough to train the reference config. - **Get it**: download the tarballs from the project page and extract so the layout matches `DATA_FORMAT.md` §6 (`calibration/`, `models/`, `YYYYMMDD-subject-NN///...`). Meshes are the YCB `textured_simple.obj` (already in the DexYCB `models/` release). - **Conversion** 🔵: obj pose `pose_y` in `labels_XXXXXX.npz` is already obj→cam OpenCV meters (no flip). depth `aligned_depth_to_color_*.png` uint16 mm → `/1000`. K from `calibration/intrinsics/_640x480.yml`. extrinsics = per-serial tag→cam from `calibration/extrinsics_*/extrinsics.yml`. mask = `seg == ycb_id`. mesh `models//textured_simple.obj` (m). ## YCB-V (train; 21 YCB objects) — subdir `ycbv/` - **Host**: BOP — https://bop.felk.cvut.cz/datasets/ (dataset "YCB-V"). - **Access**: open. **Note**: the training split usually must be fetched from BOP (`ycbv_train_real.zip`); only models + test may already be local. - **Size**: base+models ~10 GB; `train_real` ~= 60 GB. - **Get it** (BOP WebDAV / HTTPS): ```bash cd $MIGRATOR_DATA_ROOT/ycbv for f in ycbv_base ycbv_models ycbv_train_real ycbv_test_bop19; do wget https://bop.felk.cvut.cz/media/data/bop_datasets/${f}.zip && unzip ${f}.zip done ``` - **Conversion** 🔵 (BOP, native OpenCV, no flip): `R_oc=cam_R_m2c.reshape(3,3)`, `t_oc=cam_t_m2c*1e-3`. mesh `models_eval/obj_{id:06d}.ply /1000`. `depth=raw*depth_scale/1000`. K=`cam_K`. extrinsics `cam_R/t_w2c` (mm→m) world→cam. masks `mask_visib/`. ## PACE (train; 238 non-YCB objects) — subdir `pace/` - **Host**: PACE project / Hugging Face (search "PACE 6DoF pose", ~295 GB). Use the **real** split (3× RealSense rig); PACE-Sim is single-view PBR and yields no multi-view pairs. - **Access**: open (HF). License mostly MIT. - **Get it**: `huggingface-cli download --repo-type dataset --local-dir $MIGRATOR_DATA_ROOT/pace`. - **Conversion** ✅ (IoU 0.90–0.97, BOP-style): `R_oc=cam_R_m2c.reshape(3,3)`, `t_oc=cam_t_m2c/1000`. mesh `models_eval/obj_XXXXXX.ply /1000`. `depth=raw*depth_scale/1000` (real depth_scale=1.0; PBR /10000). K=`cam_K`. ⚠️ real split has NO `cam_R/t_w2c` — the 3-cam rig is stored as separate scene folders, per-view poses are self-contained (static-scene regime). ⚠️ **FILTER ARTICULATED**: drop `obj_id 545..692` (AKB-48); keep `0..544` rigid. masks `{imgid:06d}_{gtidx:06d}.png` (gtidx into `scene_gt`). ## GraspNet-1B (train + `test_novel` held-out; ~56 non-YCB of 88) — subdir `graspnet/` - **Host**: https://graspnet.net/ (download page). **License CC BY-NC** (non-commercial). - **Size**: ~170 GB. - **Get it**: register/agree on the site, download the Kinect+RealSense scene archives + `models/`, extract into `graspnet/`. - **Conversion** ✅ (IoU 0.981): `R_oc=meta['poses'][:3,:3,i]`, `t_oc=meta['poses'][:3,3,i]` (m). mesh `models/%03d/nontextured.ply` (m). ⚠️ **OBJECT-ID OFF-BY-ONE**: `cls_indexes` value `V` → mesh `models/(V-1)`. `depth=raw/1000`. K=`camK.npy` (Kinect ≠ RealSense — read per-scene). per-view poses already in cam frame (static-scene regime; no camera-pose file needed). ## HouseCat6D (TRAIN + held-out benchmark; 194 non-YCB) — subdir `housecat6d/extracted/` - **Host**: https://sites.google.com/view/housecat6d (project). Open. - **Size**: ~59 GB train (`.sqsh`/squashfs) + extracted test. - **Get it**: download from the project page; mount/extract the squashfs so `housecat6d/extracted//...` exists. - **Conversion** ✅ (IoU 0.854): `R_oc,t_oc` = label pkl `rotations[i]`, `translations[i]` — ALREADY obj→cam OpenCV meters (identity conversion). mesh `obj_models_small_size_final//.obj` (m, origin-centered; `cat`=token before `-`). `depth /1000`. K=`intrinsics.txt`. `camera_pose/NNNNNN.txt` is CAM→WORLD (invert for world→cam); train/val have NO extrinsics — use the pkl poses directly. ⚠️ masks: bg=255, use `mask != 255` (NOT `>0`). Index by `model_list`, not `class_ids`. Do NOT rescale by `scales`. ## HO-Cap (TRAIN-only pseudo-GT; 64 non-YCB) — subdir `hocap/` - **Host**: HO-Cap project (Hugging Face / project page). License CC BY 4.0. - **Size**: ~102 GB. GT is pseudo (FoundationPose+SDF); meshes BundleSDF-reconstructed. - **Get it**: download the per-subject zips + the poses/labels zips (they ship separately) and the `models/` zip; extract under `hocap/`. - **Conversion** ✅ (IoU 0.77–0.87): per-frame `label_.npz['obj_poses']` is already per-cam obj→cam (verified 1e-7) — use it directly. Otherwise `T_oc(s)=rs_RTs_inv[s]@quat_to_mat(poses_o[o,f])` with `poses_o.npy` `[qx,qy,qz,qw,tx,ty,tz]` (scipy xyzw, obj→world, world=tag_1). mesh `models//textured_mesh.obj` (m). `depth /1000` with the COLOR K. ⚠️ seg ids are LOCAL 1..4 (objects) / 5 rh / 6 lh. ## HOGraspNet (train; 30 objects, 22 YCB + 8 new) — subdir `hograspnet/` - **Host**: HOGraspNet project — **gated by a request form**. Annotations are gated; the scanned meshes are downloadable. - **Size**: ~large; ~375k states. - **Get it**: fill the access form on the project page; you receive links to the images/annotations. Meshes: `obj_scanned_models/`. - **Conversion** 🔵: `object_mat` = 4×4 obj→world (world=mas cam); `extrinsic` = (3,4) world→cam; `T_oc = M4(extrinsic) @ object_mat`; `R_oc=T_oc[:3,:3]`, `t_oc=T_oc[:3,3]/100` (cm→m). mesh `obj_scanned_models/.obj` with `p_cad = (S*V_raw)/100`, `S` = `_OBJECT_SCALE_FIXED[idx-1]` (⚠️ MANDATORY per-obj scale: wine_glass .830, small_marker .1035, spatula .671, golf_ball .430, else 1.0). `depth /1000`. K OpenCV (full-res 1920×1080; for the 640×480 crop shift `cx,cy` by the bbox). ⚠️ use the **scanned** meshes, NOT dexycb CAD (different frame). masks pseudo. ## H2O (train; 8 non-YCB objects) — subdir `h2o/` - **Host**: https://taeinkwon.com/projects/h2o/ — **gated by registration**. - **Size**: large (full H2O). We use the object-pose subset. - **Get it**: register on the project site to obtain the download credentials/links. - **Conversion** 🔵: `obj_pose_RT/{f}.txt` = class_id + 4×4 obj→cam OpenCV meters, **USED AS-IS** (⚠️ NO inversion — the "cam_to_obj" note refers to a DIFFERENT file). mesh `object/{class}.ply` (m). `depth /1000`. `cam_pose/` = cam→world (invert for world→cam). K from `cam_intrinsics.txt` (`fx fy cx cy w h`; images undistorted). Masks must be rendered from mesh+pose. Mesh folder names ≠ display names. ## DexH2R (TRAIN-only pseudo-GT; 56 non-YCB; depth on 6/18 views) — subdir `dexh2r/` - **Host**: DexH2R project — distributed via **Google Drive** (subject to Drive quota / "download quota exceeded" errors). - **Access**: open but quota-limited. **Workaround**: use `rclone` with your own Google account to copy the shared folder, which bypasses the anonymous per-file quota: ```bash rclone config # add a "gdrive" remote (your account, with the folder added to "My Drive" or shared) rclone copy gdrive:DexH2R $MIGRATOR_DATA_ROOT/dexh2r --drive-shared-with-me -P ``` Meshes + per-trial `calibration/` are small and reliable; the frame archives are the quota-blocked part. - **Conversion** 🔵 (round-trip 2.4e-16): `obj_pose.pt [T,4,4]` = `T_world<-obj` (world = hand_arm base). Kinect i: `T_oc = inv(cali0i) @ hand_arm_mesh_to_kinect_pcd_0 @ obj_pose[t]`. RealSense i(0,1): time-varying via ShadowHand FK(`qpos.pt`) @ `inv(realsense_to_forearm_i)`. mesh `object_model/.obj` (m). `depth int16→uint16 /1000`. K per-cam (Kinect 8-param distort; images undistorted). ⚠️ ONLY kinect0-3 + realsense0-1 have depth (drop 12 ZCAM). No 2D masks — project `real_obj_pcd` or threshold depth. Use per-trial `calibration/`. ## ContactPose (TRAIN-only static grasps; 25 non-YCB) — subdir `contactpose/` - **Host**: https://contactpose.cc.gatech.edu/ (project + GitHub `facebookresearch/ContactPose`). - **Size**: 2.5 TB full, but you only need a few frames per grasp (static). - **Get it**: use the official `ContactPose` python downloader (`python scripts/download_data.py ...`) to pull the RGB-D + `object_models/`. - **Conversion** ✅ (IoU 0.979): `cTo = object_pose(cam,frame)` = obj→cam OpenCV meters (`R_oc=cTo[:3,:3]`, `t_oc=cTo[:3,3]`). mesh `object_models/.ply *1e-3` (PLY is mm). `depth /1000` (Kinect v2, ~30 mm systematic bias). ⚠️ projection needs the affine `A`: `pixels = A@K@x_cam` (`cp.P`); left/right cams are PORTRAIT 540×960 while K is 960×540 landscape. extr `wTc` = cam→world (invert → cTw). ⚠️ static: sample 1–few frames per (subject,obj,intent) — 2306 grasps, NOT 2.9 M frames. masks render from mesh+cTo. --- # HELD-OUT / EVALUATION-ONLY DATASETS (never train on these) These are reserved for evaluation. In `configs/train_integrated.yaml` they carry `role: heldout` (whole-dataset OOD) or are listed in `configs/heldout.yaml` (unseen-object reservation). See `mvsam3d/data/integrated.py` for the two held-out semantics. ## HO-3D (held-out; 8 YCB objects) — subdir `ho3d_eval/` - **Host**: https://www.tugraz.at/index.php/id/2604 / codalab. Open. - **Conversion** 🔵: `R=Rodrigues(meta['objRot'])`, `t=meta['objTrans']` (m); then `R=diag(1,-1,-1)@R`, `t=diag(1,-1,-1)@t` (OpenGL→OpenCV, both). mesh dexycb `textured_simple.obj` (m). `depth=(R_ch+G_ch*256)*1.2499e-4`. K=`meta['camMat']`. eval-split masks rendered. ⚠️ exclude SB1 (train leaks to eval). ## YCB-M (held-out; YCB objects) — subdir `ycbm/` - **Host**: https://github.com/renaultjean/ycbm or the project mirror. Open. - **Conversion** 🔵: `R = quat_xyzw→R` (no flip); `t = location*1e-2` (cm→m), object in CAMERA frame. mesh `models_aligned/ycb_models_aligned_cm /100` (NOT dexycb). `depth = raw/10000`. `camera_data` = world→cam (invert → cam→world). ⚠️ handedness taken as-is (translation validated; render-verify rotation before any use). ## HOPE-Video (held-out bonus OOD; 28 non-YCB) — subdir `hope_video_repo/` - **Host**: https://github.com/swtyree/hope-dataset (HOPE). Open. - **Conversion** 🔵: `R=P[:3,:3]`, `t=P[:3,3]*1e-2` (cm→m). mesh `meshes/eval/{class}.obj /100`. `depth raw*1e-3`. extrinsics raw-as-meters world→cam (do NOT scale). K=`camera.intrinsics`. masks `masks_visib/`. --- ## Priority for expanding training beyond DexYCB From our data survey, the highest-value additions (most NEW distinct objects, our real bottleneck) are **HouseCat6D (194) + PACE (238) + GraspNet (~56)**, taking the object count from 20 → ~520 and adding ~800k clean multi-view pairs. HO-Cap, HOGraspNet, H2O, DexH2R, ContactPose add hands/interaction diversity but fewer new rigid shapes and carry pseudo-GT or gating caveats noted above.