V-JEPA 2.1 hand trajectory predictor
This model predicts one second of 3D hand motion from eight past 30 Hz RGB instants from four synchronized cameras and eight past 3D hand states. It returns 30 future states at 30 Hz. Each state contains a palm center and 21 ordered hand landmarks in world-frame metres: [30, 22, 3].
The predictor is a spline model conditioned on a frozen V-JEPA 2.1 ViT-B encoder. One visual stream uses a hand crop selected from the supplied past 3D states; four streams use full camera views. Future RGB and future labels are never model inputs. The training and inference code and complete hand guide are on the v-jepa-hand-trajectory-prediction branch.
Files and dependencies
| File | Purpose |
|---|---|
best.pt |
Epoch-29 hand predictor: weights, optimizer state, configuration, and RNG state. |
model_manifest.json |
Input/output contract, checksums, and metric summary. |
validation_kpis.json |
Full 1,142-window validation report and baselines. |
test_kpis.json |
Full 579-window held-out recording report. |
best.pt SHA-256: 0a54b0823cfb96078996210ffc1ce41b3d5014f4870935f56d2049fe5a1e1dd3.
The required frozen backbone is Meta's official V-JEPA 2.1 ViT-B checkpoint, SHA-256 848a77c33cc9e6649ed2119c9bea1e2c569bcdab9539ff3e7c02ccc2959ddf4d. Download it separately. The inference code checks its hash against best.pt. No backbone weights, source recordings, or annotations are bundled here.
Quick start
Use Python 3.12 and a compatible PyTorch environment. A CUDA GPU is recommended for the five V-JEPA streams.
git clone --branch v-jepa-hand-trajectory-prediction \
https://github.com/Movendi-ai/vjepa-trajectory-prediction.git
cd vjepa-trajectory-prediction
python -m pip install -e . huggingface_hub
mkdir -p assets
python - <<'PY'
from huggingface_hub import hf_hub_download
hf_hub_download(repo_id="KaliberAI/vjepa-hand-trajectory-prediction",
filename="best.pt", local_dir="assets")
PY
curl -fL --retry 3 \
https://dl.fbaipublicfiles.com/vjepa2/vjepa2_1_vitb_dist_vitG_384.pt \
-o assets/vjepa2_1_vitb_dist_vitG_384.pt
echo '0a54b0823cfb96078996210ffc1ce41b3d5014f4870935f56d2049fe5a1e1dd3 assets/best.pt' | sha256sum -c -
echo '848a77c33cc9e6649ed2119c9bea1e2c569bcdab9539ff3e7c02ccc2959ddf4d assets/vjepa2_1_vitb_dist_vitG_384.pt' | sha256sum -c -
For inference, provide your own synchronized RGB, observed 3D hand states, and matching calibration:
python -m vjepa_trajectory hand-predict \
--checkpoint assets/best.pt \
--backbone assets/vjepa2_1_vitb_dist_vitG_384.pt \
--input /path/to/past_hand_input.npz \
--calibration /path/to/calibration.json \
--output prediction.npz --device cuda:0
Input NPZ fields:
camera_serials[4]:45704404,46000830,54707085,56699700in that order.rgb[4,8,H,W,3]: uint8 RGB left-eye images, oldest to newest at 30 Hz; all four cameras refer to the same eight instants.history[8,22,3]: observed world-frame XYZ in metres; index 0 is the palm center, indices 1–21 are hand landmarks.history_valid[8,22]: boolean observed-joint mask; all eight palm centers must be valid.
Calibration must map the four serials to intrinsics and camera-to-world transforms in the same world frame as history. A recording-keyed file also requires --clip <recording-id>. The output NPZ contains prediction_world_m[30,22,3], future_time_s[30] (1/30 through 1 second), and crop_camera. See the inference guide for the Python HandPredictor and cached RollingHandPredictor APIs.
This release supplies a motion predictor, not a 3D hand detector or calibration for a new rig. Live use requires eight past 3D states, four synchronized RGB streams, capture timestamps, and camera calibration.
Training and evaluation
The reference training set has 2,153 windows from 48 recordings. Validation has 1,142 windows from six disjoint recordings; held-out test has 579 windows from another recording. Each window has eight past and 30 future 30 Hz instants. The 742,857-parameter predictor uses state-to-visual attention and a smooth palm/joint spline head. Epoch 29 was selected by the full-validation palm and valid-joint average displacement error.
| Split | Windows | Palm ADE | Palm +1 s FDE | Joint ADE | Joint +1 s FDE |
|---|---|---|---|---|---|
| Full validation | 1,142 | 10.02 cm | 13.70 cm | 10.93 cm | 14.55 cm |
| Held-out test | 579 | 6.21 cm | 8.24 cm | 5.85 cm | 8.04 cm |
ADE is mean Euclidean world-coordinate error over 30 future frames; FDE is error at +1 second. Joint metrics include only valid pseudo-labeled joints. These are complete-split results. On held-out test, a static-state baseline scored 5.83 cm joint ADE, slightly better than this model's 5.85 cm. The KPI files include other baselines and per-recording results.
Limitations
- Training and scoring use automatic 3D pseudo-labels, not human-verified ground truth.
- The model was trained with four fixed camera identities and 30 Hz timing. Other rigs need matching calibration and likely adaptation.
- Fast motion, occlusion, or noisy/missing 3D input joints can reduce accuracy. The held-out joint result does not improve on the static baseline.
- This checkpoint requires both past 3D states and RGB; it is not an RGB-only hand tracker.
- The model forecasts motion. It does not identify people, estimate intent, or guarantee safe robot control.
The source recordings and annotation archive are not published with this model. The repository guide documents their expected schema and the full training recipe for teams with suitable data.