Cosmos3-baseline-DROID
A post-trained Cosmos3-Nano video-and-action policy, trained on Open X-Embodiment
(DROID included) with the stock Cosmos3-Nano architecture โ the action stream runs
through the backbone's own inline action pathway (action2llm / llm2action /
action_modality_embed), with no external action expert. It is the architecture baseline
next to Cosmos3-MV-action, which replaces that pathway with a separate ActionDiT.
Model details
| Base model | nvidia/Cosmos3-Nano (Qwen3-VL-8B MoT backbone + diffusion expert) |
| Architecture | cosmos3_omni, unified_3d_mrope โ unmodified, inline action pathway (action_gen=true) |
| Parameters | 15.19 B (incl. the Qwen3-VL ViT tower) |
| Weights | EMA weights, bf16 |
| Training iteration | 250 (early checkpoint of a 5000-step schedule) |
| Warm start | stock Cosmos3-Nano |
| Action space | 64-dim zero-padded, domain-aware I/O, 64 embodiment domains |
Training
- Data โ Open X-Embodiment,
v3_full_rdt_adaptedpreset (23 dataset-weighted rows, DROID and Bridge among them) at 256p, all camera views per episode (require_all_views), action chunk 16, history latents {0,1,2}, history actions and state on, action-channel masking, 84k tokens per packed sample. - Objective โ rectified-flow video loss and action loss (action weight 10), independent action noise schedule, no diffusion forcing on the video side.
- Optimization โ LR 2e-5, warmup-cosine schedule over 5000 steps.
- Normalizer โ 3DA GAM native base-delta action statistics.
- Camera conditioning is off (OXE has no calibration).
Usage
A standard consolidated Cosmos checkpoint (config.json + sharded model*.safetensors +
checkpoint.json), the same layout nvidia/Cosmos3-Nano ships, and it loads the same way:
hf download rooty2020/Cosmos3-baseline-DROID --local-dir ./Cosmos3-baseline-DROID
torchrun --nproc_per_node=<N> -m cosmos_framework.scripts.inference \
-i inputs.json -o outputs/ --checkpoint-path ./Cosmos3-baseline-DROID
Because the action head is the stock inline pathway, action rollout needs no extra modules
beyond what the recipe's policy sampler provides; training_config.yaml carries the full
training configuration of the source run.
Provenance
Exported from a PyTorch Distributed Checkpoint (iter 250) with
python -m cosmos_framework.scripts.export_model --use-ema-weights. The ViT tower is not in
the training checkpoint and is taken from Qwen/Qwen3-VL-8B-Instruct at the revision pinned
by the Cosmos framework.
License
Derived from nvidia/Cosmos3-Nano and governed by the
NVIDIA Open Model License. Training data comes
from Open X-Embodiment and
DROID; their terms apply to the data. The usual caveats
about generated video and learned policies (no guarantee of physical accuracy, not for
safety-critical control) apply.
- Downloads last month
- 7
Model tree for rooty2020/Cosmos3-baseline-DROID
Base model
nvidia/Cosmos3-Nano