Files changed (1) hide show
  1. README.md +16 -2
README.md CHANGED
@@ -36,6 +36,18 @@ This model is ready for commercial and non-commercial use.
36
  - Cosmos3-Nano:
37
  - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
38
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  ### License
40
 
41
  This model is released under the [OpenMDW1.1](https://openmdw.ai/license/1-1/)
@@ -66,6 +78,10 @@ Cosmos3 is an Omni-modal foundation model built on a Mixture-of-Transformers (Mo
66
  **Number of trainable model parameters:**
67
 
68
  - Cosmos3-Nano: 16B
 
 
 
 
69
 
70
  ## Input/Output Specifications
71
 
@@ -184,8 +200,6 @@ Raw data from internal and external sources is transformed into training-ready d
184
 
185
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
186
 
187
- - For more information about the datasets used to train this model, please see the [Public Summary of Training Content](https://docs.nvidia.com/cosmos/latest/_downloads/e482b7114ce8dbfbb07d2d4b42cafe4e/training-content-Cosmos-3.pdf).
188
-
189
  **Data Modality and Training Data Size**
190
 
191
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |
 
36
  - Cosmos3-Nano:
37
  - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
38
 
39
+ - Cosmos3-Super:
40
+ - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
41
+
42
+ - Cosmos3-Nano-Policy-DROID:
43
+ - Given language instructions and visual observations from the DROID robot platform, generate robot action trajectories for manipulation and control tasks.
44
+
45
+ - Cosmos3-Super-Image2Video:
46
+ - Given one input image and text instructions, generate temporally coherent video sequences that are consistent with the provided visual content.
47
+
48
+ - Cosmos3-Super-Text2Image:
49
+ - Given text input, generate high-fidelity images that are consistent with the provided description.
50
+
51
  ### License
52
 
53
  This model is released under the [OpenMDW1.1](https://openmdw.ai/license/1-1/)
 
78
  **Number of trainable model parameters:**
79
 
80
  - Cosmos3-Nano: 16B
81
+ - Cosmos3-Super: 64B
82
+ - Cosmos3-Nano-Policy-DROID: 16B
83
+ - Cosmos3-Super-Image2Video: 64B
84
+ - Cosmos3-Super-Text2Image: 64B
85
 
86
  ## Input/Output Specifications
87
 
 
200
 
201
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
202
 
 
 
203
  **Data Modality and Training Data Size**
204
 
205
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |