Cosmos3-ours-RB-Y1

A Cosmos3-Nano multi-view video world model (Omni-4D) fine-tuned on RB-Y1 + WujiHand2 teleoperation with ego and exo views. Beyond video prediction it carries a 3D point-tracking head trained jointly with the backbone and explicit camera conditioning, so a rollout comes with per-frame 3D trajectories of the robot's arm and hand points.

Model details

Base model nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert)
Architecture cosmos3_omni, unified_3d_mrope, Omni-4D multi-view packing, + tracking_head, + camera_conditioner
Parameters 15.22 B (incl. the Qwen3-VL ViT tower)
Weights EMA weights, bf16 (tracking head and camera conditioner included)
Training iteration 800 (of a 4000-step schedule)
Warm start the Omni-4D stage-3 multi-view video+tracking model (0902_s3_full, iter 1068)

Training

  • Data โ€” RB-Y1 WujiHand2 teleop, 100 episodes, 2 views (ego + exo) resized 1280ร—720 โ†’ 832ร—480. Per-episode captions.
  • 3D point tracks โ€” 768 points per frame (512 on the arms, 256 on the hands) in the robot base frame, projected into both views; the exo view is the anchor (exact calibration). Visibility is "finite, in front of the camera and in bounds" only โ€” no depth ships, so self-occlusion is not modelled.
  • Cameras โ€” both cameras are static in the base frame (verified: zero pose variance on every episode), so the future camera stream is held constant.
  • Windows โ€” 4n+1 tiling plus a tail window so every frame is supervised: 741 windows (703 train / 38 val).
  • Optimization โ€” LR 1e-4, one linear schedule over 4000 steps.

Files

config.json, model-0000{1..7}-of-00007.safetensors, model.safetensors.index.json
checkpoint.json                 export provenance (use_ema_weights: true)
training_config.yaml            the full training config of the source run
dcp/iter_000000800/             the same checkpoint as a raw DCP (weights only, net + net_ema)

Usage

A standard consolidated Cosmos checkpoint, the same layout nvidia/Cosmos3-Nano ships:

hf download rooty2020/Cosmos3-ours-RB-Y1 --local-dir ./Cosmos3-ours-RB-Y1 --exclude "dcp/*"

torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
    -i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-ours-RB-Y1

The tracking head and camera conditioning are specific to Omni-4D; stock cosmos-predict without those modules loads the video backbone only.

Provenance

Exported from a PyTorch Distributed Checkpoint (iter 800) with python -m cosmos_framework.scripts.export_model --use-ema-weights. All 1069 net_ema tensors of the training checkpoint are present in the export (tracking head and camera conditioner included). The ViT tower is not in the training checkpoint and is taken from Qwen/Qwen3-VL-8B-Instruct at the revision pinned by the Cosmos framework.

License

Derived from nvidia/Cosmos3-Nano and governed by the NVIDIA Open Model License. The usual caveats about generated video (no guarantee of physical accuracy, not for safety-critical control) apply.

Downloads last month
19
Safetensors
Model size
15B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rooty2020/Cosmos3-ours-RB-Y1

Finetuned
(35)
this model