Cosmos3-MV-action

A post-trained Cosmos3-Nano multi-view video world model with an external action expert (ActionDiT), trained jointly on video and robot actions over Open X-Embodiment + DROID.

The video backbone predicts future frames; the action expert reads the backbone's per-layer K/V (one-way coupling) and denoises a single robot-action stream in parallel, so one forward gives both a video rollout and an action chunk.

Model details

Base model nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert)
Architecture cosmos3_omni, unified_3d_mrope position embeddings, + external ActionDiT
Parameters 15.16 B backbone (incl. the Qwen3-VL ViT tower) + 0.85 B action expert
Weights EMA weights โ€” backbone bf16, action expert fp32
Training iteration 9912
Warm start the stage-3 multi-view video model (Cosmos-H2R-0919, iter 2000)
Action space 64-dim zero-padded, domain-aware I/O, 64 embodiment domains

Action expert

A 36-layer, 1024-wide DiT (32 heads / 8 KV heads / head_dim 128, SwiGLU, RMSNorm, per-head RMS q/k norm, shared adaLN-zero modulation, fused self+cross attention). Each expert block cross-attends the matching backbone layer's K/V (layer_mapping=identity, cross_kv_source=gen_und, coupling=one_way, K/V not detached), and the backbone co-trains on the video loss (train_backbone=true). It was initialized by interpolating the backbone's generation tower (init=interp_from_gen). The backbone carries no inline action pathway (action_gen=false in config.json) โ€” the expert replaces it โ€” and camera conditioning is off, since OXE has no calibration.

Training

  • Data โ€” Open X-Embodiment, v3_full_rdt_adapted preset (23 dataset-weighted rows, DROID and Bridge among them), 2 views (third-person + wrist) at 256p, action chunk 16, history latents {0,1,2}, 24k tokens per packed sample.
  • Objective โ€” rectified-flow video loss (ร—10) plus a rectified-flow action loss (ร—10), with an independent action noise schedule (the video side uses diffusion forcing).
  • Normalizer โ€” 3DA GAM native base-delta action statistics.

Files

config.json                     backbone (Cosmos3OmniModel) config
model-0000{1..7}-of-00007.safetensors, model.safetensors.index.json
checkpoint.json                 export provenance (use_ema_weights: true)
action_expert/action_dit.safetensors   ActionDiT weights, keys action_dit.*
action_expert/config.json       resolved ActionExpertConfig + provenance
training_config.yaml            the full training config of the source run

The backbone half is a standard consolidated Cosmos checkpoint and loads with the stock Cosmos inference entry point:

hf download rooty2020/Cosmos3-MV-action --local-dir ./Cosmos3-MV-action

torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
    -i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-MV-action

That path gives video only. To run the policy, build an ActionExpertMoTModel with model.config.action_expert set from action_expert/config.json, load the backbone safetensors into net, and load action_dit.safetensors into net.action_dit (keys already carry the action_dit. prefix).

Provenance

Exported from a PyTorch Distributed Checkpoint (iter 9912) with cosmos_framework.scripts.export_action_expert_model: the backbone through the stock export_model path with action_gen=False, and net_ema.action_dit.* pulled straight out of the same DCP. The ViT tower is not in the training checkpoint and is taken from Qwen/Qwen3-VL-8B-Instruct at the revision pinned by the Cosmos framework.

License

Derived from nvidia/Cosmos3-Nano and governed by the NVIDIA Open Model License. Training data comes from Open X-Embodiment and DROID; their terms apply to the data. The usual caveats about generated video and learned policies (no guarantee of physical accuracy, not for safety-critical control) apply.

Downloads last month
16
Safetensors
Model size
15B params
Tensor type
BF16
ยท
Video Preview
loading

Model tree for rooty2020/Cosmos3-MV-action

Finetuned
(31)
this model