Instructions to use rooty2020/Cosmos3-ours-RB-Y1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use rooty2020/Cosmos3-ours-RB-Y1 with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Cosmos3-ours-RB-Y1
A Cosmos3-Nano multi-view video world model (Omni-4D) fine-tuned on RB-Y1 + WujiHand2 teleoperation with ego and exo views. Beyond video prediction it carries a 3D point-tracking head trained jointly with the backbone and explicit camera conditioning, so a rollout comes with per-frame 3D trajectories of the robot's arm and hand points.
Model details
| Base model | nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert) |
| Architecture | cosmos3_omni, unified_3d_mrope, Omni-4D multi-view packing, + tracking_head, + camera_conditioner |
| Parameters | 15.22 B (incl. the Qwen3-VL ViT tower) |
| Weights | EMA weights, bf16 (tracking head and camera conditioner included) |
| Training iteration | 800 (of a 4000-step schedule) |
| Warm start | the Omni-4D stage-3 multi-view video+tracking model (0902_s3_full, iter 1068) |
Training
- Data โ RB-Y1 WujiHand2 teleop, 100 episodes, 2 views (ego + exo) resized 1280ร720 โ 832ร480. Per-episode captions.
- 3D point tracks โ 768 points per frame (512 on the arms, 256 on the hands) in the robot base frame, projected into both views; the exo view is the anchor (exact calibration). Visibility is "finite, in front of the camera and in bounds" only โ no depth ships, so self-occlusion is not modelled.
- Cameras โ both cameras are static in the base frame (verified: zero pose variance on every episode), so the future camera stream is held constant.
- Windows โ 4n+1 tiling plus a tail window so every frame is supervised: 741 windows (703 train / 38 val).
- Optimization โ LR 1e-4, one linear schedule over 4000 steps.
Files
config.json, model-0000{1..7}-of-00007.safetensors, model.safetensors.index.json
checkpoint.json export provenance (use_ema_weights: true)
training_config.yaml the full training config of the source run
dcp/iter_000000800/ the same checkpoint as a raw DCP (weights only, net + net_ema)
Usage
A standard consolidated Cosmos checkpoint, the same layout nvidia/Cosmos3-Nano ships:
hf download rooty2020/Cosmos3-ours-RB-Y1 --local-dir ./Cosmos3-ours-RB-Y1 --exclude "dcp/*"
torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
-i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-ours-RB-Y1
The tracking head and camera conditioning are specific to Omni-4D; stock cosmos-predict
without those modules loads the video backbone only.
Provenance
Exported from a PyTorch Distributed Checkpoint (iter 800) with
python -m cosmos_framework.scripts.export_model --use-ema-weights. All 1069 net_ema tensors
of the training checkpoint are present in the export (tracking head and camera conditioner
included). The ViT tower is not in the training checkpoint and is taken from
Qwen/Qwen3-VL-8B-Instruct at the revision pinned by the Cosmos framework.
License
Derived from nvidia/Cosmos3-Nano and governed by the
NVIDIA Open Model License. The usual caveats
about generated video (no guarantee of physical accuracy, not for safety-critical control)
apply.
- Downloads last month
- 19
Model tree for rooty2020/Cosmos3-ours-RB-Y1
Base model
nvidia/Cosmos3-Nano