Download migrator/docs/DOWNLOAD_GUIDE.md from Ronaldo-GOAT/bert_simpson: direct link, hf CLI and curl.
- Browser
- Download file 12.2 kB
-
https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/migrator/docs/DOWNLOAD_GUIDE.md
- Command line
-
hf download hf://Ronaldo-GOAT/bert_simpson/migrator/docs/DOWNLOAD_GUIDE.md
-
curl -L -o DOWNLOAD_GUIDE.md https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/migrator/docs/DOWNLOAD_GUIDE.md
DOWNLOAD_GUIDE — per-dataset acquisition + conversion
How to obtain each dataset we integrated and convert it into the unified contract
(DATA_FORMAT.md). For each dataset: host/URL, access (open vs gated),
approx size, how to get it, and the coordinate conversion to our
contract (condensed from the verified
code/mv-sam3d-for-6d-v2/docs/dataset_convention_spec.md — that file is the
authority if anything here is ambiguous).
Put everything under one root and point the config at it:
export MIGRATOR_DATA_ROOT=/path/to/datasets # each dataset in its own subdir
After downloading a dataset, build its canonical SS latents once (see
DATA_FORMAT.md §4) before training on it.
Legend for conversion status (from the spec): ✅ = empirically reprojection-IoU verified; 🔵 = code/cross-check verified.
Notation: R_oc,t_oc = object(CAD)→cam OpenCV meters (x_cam=R_oc·p_cad+t_oc);
p_cad = mesh verts (m). Adapters return raw OpenCV meters — never pre-apply the
D=diag(-1,-1,1) flip or the canonical M (the base loader does that).
TRAINING DATASETS
DexYCB (reference; 20 YCB objects) — subdir dexycb/
- Host: https://dex-ycb.github.io/ (NVIDIA). Open, registration-free download links on the project page ("Download").
- Size: ~1.4 TB full; the s0/s1 subset + models is enough to train the reference config.
- Get it: download the tarballs from the project page and extract so the
layout matches
DATA_FORMAT.md§6 (calibration/,models/,YYYYMMDD-subject-NN/<seq>/<serial>/...). Meshes are the YCBtextured_simple.obj(already in the DexYCBmodels/release). - Conversion 🔵: obj pose
pose_yinlabels_XXXXXX.npzis already obj→cam OpenCV meters (no flip). depthaligned_depth_to_color_*.pnguint16 mm →/1000. K fromcalibration/intrinsics/<serial>_640x480.yml. extrinsics = per-serial tag→cam fromcalibration/extrinsics_*/extrinsics.yml. mask =seg == ycb_id. meshmodels/<obj>/textured_simple.obj(m).
YCB-V (train; 21 YCB objects) — subdir ycbv/
- Host: BOP — https://bop.felk.cvut.cz/datasets/ (dataset "YCB-V").
- Access: open. Note: the training split usually must be fetched from BOP
(
ycbv_train_real.zip); only models + test may already be local. - Size: base+models ~10 GB;
train_real~= 60 GB. - Get it (BOP WebDAV / HTTPS):
cd $MIGRATOR_DATA_ROOT/ycbv for f in ycbv_base ycbv_models ycbv_train_real ycbv_test_bop19; do wget https://bop.felk.cvut.cz/media/data/bop_datasets/${f}.zip && unzip ${f}.zip done - Conversion 🔵 (BOP, native OpenCV, no flip):
R_oc=cam_R_m2c.reshape(3,3),t_oc=cam_t_m2c*1e-3. meshmodels_eval/obj_{id:06d}.ply /1000.depth=raw*depth_scale/1000. K=cam_K. extrinsicscam_R/t_w2c(mm→m) world→cam. masksmask_visib/.
PACE (train; 238 non-YCB objects) — subdir pace/
- Host: PACE project / Hugging Face (search "PACE 6DoF pose", ~295 GB). Use the real split (3× RealSense rig); PACE-Sim is single-view PBR and yields no multi-view pairs.
- Access: open (HF). License mostly MIT.
- Get it:
huggingface-cli download <pace-repo> --repo-type dataset --local-dir $MIGRATOR_DATA_ROOT/pace. - Conversion ✅ (IoU 0.90–0.97, BOP-style):
R_oc=cam_R_m2c.reshape(3,3),t_oc=cam_t_m2c/1000. meshmodels_eval/obj_XXXXXX.ply /1000.depth=raw*depth_scale/1000(real depth_scale=1.0; PBR /10000). K=cam_K. ⚠️ real split has NOcam_R/t_w2c— the 3-cam rig is stored as separate scene folders, per-view poses are self-contained (static-scene regime). ⚠️ FILTER ARTICULATED: dropobj_id 545..692(AKB-48); keep0..544rigid. masks{imgid:06d}_{gtidx:06d}.png(gtidx intoscene_gt).
GraspNet-1B (train + test_novel held-out; ~56 non-YCB of 88) — subdir graspnet/
- Host: https://graspnet.net/ (download page). License CC BY-NC (non-commercial).
- Size: ~170 GB.
- Get it: register/agree on the site, download the Kinect+RealSense scene
archives +
models/, extract intograspnet/. - Conversion ✅ (IoU 0.981):
R_oc=meta['poses'][:3,:3,i],t_oc=meta['poses'][:3,3,i](m). meshmodels/%03d/nontextured.ply(m). ⚠️ OBJECT-ID OFF-BY-ONE:cls_indexesvalueV→ meshmodels/(V-1).depth=raw/1000. K=camK.npy(Kinect ≠ RealSense — read per-scene). per-view poses already in cam frame (static-scene regime; no camera-pose file needed).
HouseCat6D (TRAIN + held-out benchmark; 194 non-YCB) — subdir housecat6d/extracted/
- Host: https://sites.google.com/view/housecat6d (project). Open.
- Size: ~59 GB train (
.sqsh/squashfs) + extracted test. - Get it: download from the project page; mount/extract the squashfs so
housecat6d/extracted/<scene>/...exists. - Conversion ✅ (IoU 0.854):
R_oc,t_oc= label pklrotations[i],translations[i]— ALREADY obj→cam OpenCV meters (identity conversion). meshobj_models_small_size_final/<cat>/<model>.obj(m, origin-centered;cat=token before-).depth /1000. K=intrinsics.txt.camera_pose/NNNNNN.txtis CAM→WORLD (invert for world→cam); train/val have NO extrinsics — use the pkl poses directly. ⚠️ masks: bg=255, usemask != 255(NOT>0). Index bymodel_list, notclass_ids. Do NOT rescale byscales.
HO-Cap (TRAIN-only pseudo-GT; 64 non-YCB) — subdir hocap/
- Host: HO-Cap project (Hugging Face / project page). License CC BY 4.0.
- Size: ~102 GB. GT is pseudo (FoundationPose+SDF); meshes BundleSDF-reconstructed.
- Get it: download the per-subject zips + the poses/labels zips (they ship
separately) and the
models/zip; extract underhocap/. - Conversion ✅ (IoU 0.77–0.87): per-frame
label_<frame>.npz['obj_poses']is already per-cam obj→cam (verified 1e-7) — use it directly. OtherwiseT_oc(s)=rs_RTs_inv[s]@quat_to_mat(poses_o[o,f])withposes_o.npy[qx,qy,qz,qw,tx,ty,tz](scipy xyzw, obj→world, world=tag_1). meshmodels/<id>/textured_mesh.obj(m).depth /1000with the COLOR K. ⚠️ seg ids are LOCAL 1..4 (objects) / 5 rh / 6 lh.
HOGraspNet (train; 30 objects, 22 YCB + 8 new) — subdir hograspnet/
- Host: HOGraspNet project — gated by a request form. Annotations are gated; the scanned meshes are downloadable.
- Size: ~large; ~375k states.
- Get it: fill the access form on the project page; you receive links to the
images/annotations. Meshes:
obj_scanned_models/. - Conversion 🔵:
object_mat= 4×4 obj→world (world=mas cam);extrinsic= (3,4) world→cam;T_oc = M4(extrinsic) @ object_mat;R_oc=T_oc[:3,:3],t_oc=T_oc[:3,3]/100(cm→m). meshobj_scanned_models/<name>.objwithp_cad = (S*V_raw)/100,S=_OBJECT_SCALE_FIXED[idx-1](⚠️ MANDATORY per-obj scale: wine_glass .830, small_marker .1035, spatula .671, golf_ball .430, else 1.0).depth /1000. K OpenCV (full-res 1920×1080; for the 640×480 crop shiftcx,cyby the bbox). ⚠️ use the scanned meshes, NOT dexycb CAD (different frame). masks pseudo.
H2O (train; 8 non-YCB objects) — subdir h2o/
- Host: https://taeinkwon.com/projects/h2o/ — gated by registration.
- Size: large (full H2O). We use the object-pose subset.
- Get it: register on the project site to obtain the download credentials/links.
- Conversion 🔵:
obj_pose_RT/{f}.txt= class_id + 4×4 obj→cam OpenCV meters, USED AS-IS (⚠️ NO inversion — the "cam_to_obj" note refers to a DIFFERENT file). meshobject/{class}.ply(m).depth /1000.cam_pose/= cam→world (invert for world→cam). K fromcam_intrinsics.txt(fx fy cx cy w h; images undistorted). Masks must be rendered from mesh+pose. Mesh folder names ≠ display names.
DexH2R (TRAIN-only pseudo-GT; 56 non-YCB; depth on 6/18 views) — subdir dexh2r/
- Host: DexH2R project — distributed via Google Drive (subject to Drive quota / "download quota exceeded" errors).
- Access: open but quota-limited. Workaround: use
rclonewith your own Google account to copy the shared folder, which bypasses the anonymous per-file quota:Meshes + per-trialrclone config # add a "gdrive" remote (your account, with the folder added to "My Drive" or shared) rclone copy gdrive:DexH2R $MIGRATOR_DATA_ROOT/dexh2r --drive-shared-with-me -Pcalibration/are small and reliable; the frame archives are the quota-blocked part. - Conversion 🔵 (round-trip 2.4e-16):
obj_pose.pt [T,4,4]=T_world<-obj(world = hand_arm base). Kinect i:T_oc = inv(cali0i) @ hand_arm_mesh_to_kinect_pcd_0 @ obj_pose[t]. RealSense i(0,1): time-varying via ShadowHand FK(qpos.pt) @inv(realsense_to_forearm_i). meshobject_model/<name>.obj(m).depth int16→uint16 /1000. K per-cam (Kinect 8-param distort; images undistorted). ⚠️ ONLY kinect0-3 + realsense0-1 have depth (drop 12 ZCAM). No 2D masks — projectreal_obj_pcdor threshold depth. Use per-trialcalibration/.
ContactPose (TRAIN-only static grasps; 25 non-YCB) — subdir contactpose/
- Host: https://contactpose.cc.gatech.edu/ (project + GitHub
facebookresearch/ContactPose). - Size: 2.5 TB full, but you only need a few frames per grasp (static).
- Get it: use the official
ContactPosepython downloader (python scripts/download_data.py ...) to pull the RGB-D +object_models/. - Conversion ✅ (IoU 0.979):
cTo = object_pose(cam,frame)= obj→cam OpenCV meters (R_oc=cTo[:3,:3],t_oc=cTo[:3,3]). meshobject_models/<obj>.ply *1e-3(PLY is mm).depth /1000(Kinect v2, ~30 mm systematic bias). ⚠️ projection needs the affineA:pixels = A@K@x_cam(cp.P); left/right cams are PORTRAIT 540×960 while K is 960×540 landscape. extrwTc= cam→world (invert → cTw). ⚠️ static: sample 1–few frames per (subject,obj,intent) — 2306 grasps, NOT 2.9 M frames. masks render from mesh+cTo.
HELD-OUT / EVALUATION-ONLY DATASETS (never train on these)
These are reserved for evaluation. In configs/train_integrated.yaml they carry
role: heldout (whole-dataset OOD) or are listed in configs/heldout.yaml
(unseen-object reservation). See mvsam3d/data/integrated.py for the two
held-out semantics.
HO-3D (held-out; 8 YCB objects) — subdir ho3d_eval/
- Host: https://www.tugraz.at/index.php/id/2604 / codalab. Open.
- Conversion 🔵:
R=Rodrigues(meta['objRot']),t=meta['objTrans'](m); thenR=diag(1,-1,-1)@R,t=diag(1,-1,-1)@t(OpenGL→OpenCV, both). mesh dexycbtextured_simple.obj(m).depth=(R_ch+G_ch*256)*1.2499e-4. K=meta['camMat']. eval-split masks rendered. ⚠️ exclude SB1 (train leaks to eval).
YCB-M (held-out; YCB objects) — subdir ycbm/
- Host: https://github.com/renaultjean/ycbm or the project mirror. Open.
- Conversion 🔵:
R = quat_xyzw→R(no flip);t = location*1e-2(cm→m), object in CAMERA frame. meshmodels_aligned/ycb_models_aligned_cm /100(NOT dexycb).depth = raw/10000.camera_data= world→cam (invert → cam→world). ⚠️ handedness taken as-is (translation validated; render-verify rotation before any use).
HOPE-Video (held-out bonus OOD; 28 non-YCB) — subdir hope_video_repo/
- Host: https://github.com/swtyree/hope-dataset (HOPE). Open.
- Conversion 🔵:
R=P[:3,:3],t=P[:3,3]*1e-2(cm→m). meshmeshes/eval/{class}.obj /100.depth raw*1e-3. extrinsics raw-as-meters world→cam (do NOT scale). K=camera.intrinsics. masksmasks_visib/.
Priority for expanding training beyond DexYCB
From our data survey, the highest-value additions (most NEW distinct objects, our real bottleneck) are HouseCat6D (194) + PACE (238) + GraspNet (~56), taking the object count from 20 → ~520 and adding ~800k clean multi-view pairs. HO-Cap, HOGraspNet, H2O, DexH2R, ContactPose add hands/interaction diversity but fewer new rigid shapes and carry pseudo-GT or gating caveats noted above.