Cosmos3-MV-action
A post-trained Cosmos3-Nano multi-view video world model with an external action expert (ActionDiT), trained jointly on video and robot actions over Open X-Embodiment + DROID.
The video backbone predicts future frames; the action expert reads the backbone's per-layer K/V (one-way coupling) and denoises a single robot-action stream in parallel, so one forward gives both a video rollout and an action chunk.
Model details
| Base model | nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert) |
| Architecture | cosmos3_omni, unified_3d_mrope position embeddings, + external ActionDiT |
| Parameters | 15.16 B backbone (incl. the Qwen3-VL ViT tower) + 0.85 B action expert |
| Weights | EMA weights โ backbone bf16, action expert fp32 |
| Training iteration | 9912 |
| Warm start | the stage-3 multi-view video model (Cosmos-H2R-0919, iter 2000) |
| Action space | 64-dim zero-padded, domain-aware I/O, 64 embodiment domains |
Action expert
A 36-layer, 1024-wide DiT (32 heads / 8 KV heads / head_dim 128, SwiGLU, RMSNorm,
per-head RMS q/k norm, shared adaLN-zero modulation, fused self+cross attention).
Each expert block cross-attends the matching backbone layer's K/V (layer_mapping=identity,
cross_kv_source=gen_und, coupling=one_way, K/V not detached), and the backbone co-trains
on the video loss (train_backbone=true). It was initialized by interpolating the backbone's
generation tower (init=interp_from_gen). The backbone carries no inline action pathway
(action_gen=false in config.json) โ the expert replaces it โ and camera conditioning is
off, since OXE has no calibration.
Training
- Data โ Open X-Embodiment,
v3_full_rdt_adaptedpreset (23 dataset-weighted rows, DROID and Bridge among them), 2 views (third-person + wrist) at 256p, action chunk 16, history latents {0,1,2}, 24k tokens per packed sample. - Objective โ rectified-flow video loss (ร10) plus a rectified-flow action loss (ร10), with an independent action noise schedule (the video side uses diffusion forcing).
- Normalizer โ 3DA GAM native base-delta action statistics.
Files
config.json backbone (Cosmos3OmniModel) config
model-0000{1..7}-of-00007.safetensors, model.safetensors.index.json
checkpoint.json export provenance (use_ema_weights: true)
action_expert/action_dit.safetensors ActionDiT weights, keys action_dit.*
action_expert/config.json resolved ActionExpertConfig + provenance
training_config.yaml the full training config of the source run
The backbone half is a standard consolidated Cosmos checkpoint and loads with the stock Cosmos inference entry point:
hf download rooty2020/Cosmos3-MV-action --local-dir ./Cosmos3-MV-action
torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
-i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-MV-action
That path gives video only. To run the policy, build an ActionExpertMoTModel with
model.config.action_expert set from action_expert/config.json, load the backbone
safetensors into net, and load action_dit.safetensors into net.action_dit
(keys already carry the action_dit. prefix).
Provenance
Exported from a PyTorch Distributed Checkpoint (iter 9912) with
cosmos_framework.scripts.export_action_expert_model: the backbone through the stock
export_model path with action_gen=False, and net_ema.action_dit.* pulled straight out of
the same DCP. The ViT tower is not in the training checkpoint and is taken from
Qwen/Qwen3-VL-8B-Instruct at the revision pinned by the Cosmos framework.
License
Derived from nvidia/Cosmos3-Nano and governed by the
NVIDIA Open Model License. Training data comes
from Open X-Embodiment and
DROID; their terms apply to the data. The usual caveats
about generated video and learned policies (no guarantee of physical accuracy, not for
safety-critical control) apply.
- Downloads last month
- 16
Model tree for rooty2020/Cosmos3-MV-action
Base model
nvidia/Cosmos3-Nano