Add SGLang serving instructions

#10
README.md CHANGED
@@ -33,9 +33,21 @@ This model is ready for commercial and non-commercial use.
33
  **Model Developer:** NVIDIA
34
 
35
  ### Model Versions
 
 
 
36
  - Cosmos3-Super:
37
  - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
38
 
 
 
 
 
 
 
 
 
 
39
  ### License
40
 
41
  This model is released under the [OpenMDW1.1](https://openmdw.ai/license/1-1/)
@@ -65,7 +77,11 @@ Cosmos3 is an Omni-modal foundation model built on a Mixture-of-Transformers (Mo
65
 
66
  **Number of trainable model parameters:**
67
 
 
68
  - Cosmos3-Super: 64B
 
 
 
69
 
70
  ## Input/Output Specifications
71
 
@@ -153,7 +169,6 @@ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated sys
153
  - [PyTorch](https://github.com/nvidia/cosmos3)
154
  - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
155
  - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
156
- - [SGLang](https://github.com/sgl-project/sglang)
157
 
158
  **Supported Hardware Microarchitecture Compatibility:**
159
 
@@ -184,8 +199,6 @@ Raw data from internal and external sources is transformed into training-ready d
184
 
185
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
186
 
187
- - For more information about the datasets used to train this model, please see the [Public Summary of Training Content](https://docs.nvidia.com/cosmos/latest/_downloads/e482b7114ce8dbfbb07d2d4b42cafe4e/training-content-Cosmos-3.pdf).
188
-
189
  **Data Modality and Training Data Size**
190
 
191
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |
@@ -843,6 +856,67 @@ Example output from the command above:
843
  4. Place the flower into the red bottle.
844
  ```
845
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
846
  ### Diffusers
847
 
848
  #### Container
@@ -910,71 +984,6 @@ Example output:
910
 
911
  <video controls width="1280" height="720" src="https://huggingface.co/nvidia/Cosmos3-Super/resolve/main/assets/example_t2v_diffusers_output.mp4"></video>
912
 
913
- ### SGLang
914
-
915
- [SGLang Diffusion](https://docs.sglang.io/docs/sglang-diffusion/index) can serve `nvidia/Cosmos3-Super` through OpenAI-compatible image and video generation endpoints. Install SGLang from the main branch with diffusion dependencies, then start the server:
916
-
917
- ```bash
918
- git clone --branch main https://github.com/sgl-project/sglang.git
919
- cd sglang
920
- pip install -e "python[diffusion]"
921
- pip install "cosmos-guardrail==0.3.1"
922
-
923
- sglang serve \
924
- --model-path nvidia/Cosmos3-Super \
925
- --num-gpus 4
926
- ```
927
-
928
- Cosmos 3 support in SGLang Diffusion currently requires the SGLang main branch. Switch to a stable SGLang release once Cosmos 3 support is included there.
929
-
930
- For the video-specialized checkpoint:
931
-
932
- ```bash
933
- sglang serve \
934
- --model-path nvidia/Cosmos3-Super-Image2Video \
935
- --num-gpus 4
936
- ```
937
-
938
- Supported SGLang endpoints:
939
-
940
- | Mode | Endpoint | Notes |
941
- | --- | --- | --- |
942
- | Text to image | `POST /v1/images/generations` | Returns base64 image data by default |
943
- | Text to video | `POST /v1/videos` | Creates an async job; poll `GET /v1/videos/{id}` and download `/content` |
944
- | Image to video | `POST /v1/videos` | Upload the conditioning image with `input_reference` |
945
-
946
- Example text-to-video request:
947
-
948
- ```bash
949
- job_id=$(curl -sS -X POST http://localhost:30000/v1/videos \
950
- --form-string "prompt=A small warehouse robot moves a blue box across a clean floor." \
951
- --form-string "negative_prompt=blurry, distorted, low quality" \
952
- --form-string "size=1280x720" \
953
- --form-string "num_frames=81" \
954
- --form-string "fps=24" \
955
- --form-string "num_inference_steps=35" \
956
- --form-string "guidance_scale=4.0" \
957
- --form-string "flow_shift=10.0" \
958
- --form-string "seed=42" \
959
- --form-string 'extra_params={"guardrails":true,"use_resolution_template":false,"use_duration_template":false}' \
960
- | python -c 'import json, sys; print(json.load(sys.stdin)["id"])')
961
-
962
- while true; do
963
- status=$(curl -sS "http://localhost:30000/v1/videos/${job_id}" \
964
- | python -c 'import json, sys; print(json.load(sys.stdin)["status"])')
965
- [ "$status" = "completed" ] && break
966
- [ "$status" = "failed" ] && exit 1
967
- sleep 1
968
- done
969
-
970
- curl -sS -L "http://localhost:30000/v1/videos/${job_id}/content" \
971
- -o cosmos3_super_t2v_output.mp4
972
- ```
973
-
974
- Video-to-video, video-with-sound, and action generation are not supported by SGLang yet.
975
-
976
- For complete serving instructions and request examples, see the [Cosmos3 SGLang cookbook](https://docs.sglang.io/cookbook/diffusion/Cosmos/Cosmos3).
977
-
978
  ## Limitations
979
 
980
  Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
@@ -983,7 +992,7 @@ Cosmos3 outputs should not be treated as physically accurate simulation, reliabl
983
 
984
  ## Inference
985
 
986
- **Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers), [SGLang](https://github.com/sgl-project/sglang), [SGLang Diffusion](https://docs.sglang.io/docs/sglang-diffusion/index)
987
 
988
  **Test Hardware:** GB200 and H100
989
 
 
33
  **Model Developer:** NVIDIA
34
 
35
  ### Model Versions
36
+ - Cosmos3-Nano:
37
+ - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
38
+
39
  - Cosmos3-Super:
40
  - Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
41
 
42
+ - Cosmos3-Nano-Policy-DROID:
43
+ - Given language instructions and visual observations from the DROID robot platform, generate robot action trajectories for manipulation and control tasks.
44
+
45
+ - Cosmos3-Super-Image2Video:
46
+ - Given one input image and text instructions, generate temporally coherent video sequences that are consistent with the provided visual content.
47
+
48
+ - Cosmos3-Super-Text2Image:
49
+ - Given text input, generate high-fidelity images that are consistent with the provided description.
50
+
51
  ### License
52
 
53
  This model is released under the [OpenMDW1.1](https://openmdw.ai/license/1-1/)
 
77
 
78
  **Number of trainable model parameters:**
79
 
80
+ - Cosmos3-Nano: 16B
81
  - Cosmos3-Super: 64B
82
+ - Cosmos3-Nano-Policy-DROID: 16B
83
+ - Cosmos3-Super-Image2Video: 64B
84
+ - Cosmos3-Super-Text2Image: 64B
85
 
86
  ## Input/Output Specifications
87
 
 
169
  - [PyTorch](https://github.com/nvidia/cosmos3)
170
  - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
171
  - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
 
172
 
173
  **Supported Hardware Microarchitecture Compatibility:**
174
 
 
199
 
200
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
201
 
 
 
202
  **Data Modality and Training Data Size**
203
 
204
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |
 
856
  4. Place the flower into the red bottle.
857
  ```
858
 
859
+ ### SGLang
860
+
861
+ SGLang-Diffusion can serve `nvidia/Cosmos3-Super` through OpenAI-compatible image and video endpoints. Install SGLang from source with diffusion dependencies, then start the server:
862
+
863
+ ```bash
864
+ git clone https://github.com/sgl-project/sglang.git
865
+ cd sglang
866
+ pip install -e "python[diffusion]"
867
+ pip install "cosmos-guardrail==0.3.1"
868
+
869
+ sglang serve \
870
+ --model-path nvidia/Cosmos3-Super \
871
+ --num-gpus 4
872
+ ```
873
+
874
+ For the video-specialized checkpoint:
875
+
876
+ ```bash
877
+ sglang serve \
878
+ --model-path nvidia/Cosmos3-Super-Image2Video \
879
+ --num-gpus 4
880
+ ```
881
+
882
+ Supported SGLang endpoints:
883
+
884
+ | Mode | Endpoint | Notes |
885
+ | --- | --- | --- |
886
+ | Text to image | `POST /v1/images/generations` | Returns base64 image data by default |
887
+ | Text to video | `POST /v1/videos` | Creates an async job; poll `GET /v1/videos/{id}` and download `/content` |
888
+ | Image to video | `POST /v1/videos` | Upload the conditioning image with `input_reference` |
889
+
890
+ Example text-to-video request:
891
+
892
+ ```bash
893
+ job_id=$(curl -sS -X POST http://localhost:30000/v1/videos \
894
+ --form-string "prompt=A small warehouse robot moves a blue box across a clean floor." \
895
+ --form-string "negative_prompt=blurry, distorted, low quality" \
896
+ --form-string "size=1280x720" \
897
+ --form-string "num_frames=81" \
898
+ --form-string "fps=24" \
899
+ --form-string "num_inference_steps=35" \
900
+ --form-string "guidance_scale=4.0" \
901
+ --form-string "flow_shift=10.0" \
902
+ --form-string "seed=42" \
903
+ --form-string 'extra_params={"guardrails":true,"use_resolution_template":false,"use_duration_template":false}' \
904
+ | python -c 'import json, sys; print(json.load(sys.stdin)["id"])')
905
+
906
+ while true; do
907
+ status=$(curl -sS "http://localhost:30000/v1/videos/${job_id}" \
908
+ | python -c 'import json, sys; print(json.load(sys.stdin)["status"])')
909
+ [ "$status" = "completed" ] && break
910
+ [ "$status" = "failed" ] && exit 1
911
+ sleep 1
912
+ done
913
+
914
+ curl -sS -L "http://localhost:30000/v1/videos/${job_id}/content" \
915
+ -o cosmos3_super_t2v_output.mp4
916
+ ```
917
+
918
+ Video-to-video, video-with-sound, and action generation are not supported by SGLang yet.
919
+
920
  ### Diffusers
921
 
922
  #### Container
 
984
 
985
  <video controls width="1280" height="720" src="https://huggingface.co/nvidia/Cosmos3-Super/resolve/main/assets/example_t2v_diffusers_output.mp4"></video>
986
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
987
  ## Limitations
988
 
989
  Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
 
992
 
993
  ## Inference
994
 
995
+ **Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers)
996
 
997
  **Test Hardware:** GB200 and H100
998
 
modular_model_index.json DELETED
@@ -1,75 +0,0 @@
1
- {
2
- "_blocks_class_name": "Cosmos3OmniBlocks",
3
- "_class_name": "Cosmos3OmniModularPipeline",
4
- "_diffusers_version": "0.39.0.dev0",
5
- "text_tokenizer": [
6
- "transformers",
7
- "Qwen2TokenizerFast",
8
- {
9
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Super",
10
- "revision": null,
11
- "subfolder": "text_tokenizer",
12
- "type_hint": [
13
- "transformers",
14
- "Qwen2TokenizerFast"
15
- ],
16
- "variant": null
17
- }
18
- ],
19
- "vae": [
20
- "diffusers",
21
- "AutoencoderKLWan",
22
- {
23
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Super",
24
- "revision": null,
25
- "subfolder": "vae",
26
- "type_hint": [
27
- "diffusers",
28
- "AutoencoderKLWan"
29
- ],
30
- "variant": null
31
- }
32
- ],
33
- "transformer": [
34
- "diffusers",
35
- "Cosmos3OmniTransformer",
36
- {
37
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Super",
38
- "revision": null,
39
- "subfolder": "transformer",
40
- "type_hint": [
41
- "diffusers",
42
- "Cosmos3OmniTransformer"
43
- ],
44
- "variant": null
45
- }
46
- ],
47
- "scheduler": [
48
- "diffusers",
49
- "UniPCMultistepScheduler",
50
- {
51
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Super",
52
- "revision": null,
53
- "subfolder": "scheduler",
54
- "type_hint": [
55
- "diffusers",
56
- "UniPCMultistepScheduler"
57
- ],
58
- "variant": null
59
- }
60
- ],
61
- "sound_tokenizer": [
62
- "diffusers",
63
- "Cosmos3AVAEAudioTokenizer",
64
- {
65
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Super",
66
- "revision": null,
67
- "subfolder": "sound_tokenizer",
68
- "type_hint": [
69
- "diffusers",
70
- "Cosmos3AVAEAudioTokenizer"
71
- ],
72
- "variant": null
73
- }
74
- ]
75
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
sound_tokenizer.ckpt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6daeb68a219f3e86c0918f616d78b9ebf073f3d700df63ff1c02d214c081d72d
3
+ size 1985246007
sound_tokenizer.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "autoencoder_v2",
3
+ "sampling_rate": 48000,
4
+ "stereo": true,
5
+ "use_wav_as_input": true,
6
+ "normalize_volume": true,
7
+ "hop_size": 1920,
8
+ "input_channels": 1,
9
+ "enc_type": "spec_convnext",
10
+ "enc_dim": 192,
11
+ "enc_intermediate_dim": 768,
12
+ "enc_num_layers": 12,
13
+ "enc_num_blocks": 2,
14
+ "enc_n_fft": 64,
15
+ "enc_hop_length": 16,
16
+ "enc_latent_dim": 128,
17
+ "enc_c_mults": [1, 2, 4],
18
+ "enc_strides": [4, 5, 6],
19
+ "enc_identity_init": false,
20
+ "enc_use_snake": true,
21
+ "dec_type": "oobleck",
22
+ "dec_dim": 320,
23
+ "dec_c_mults": [1, 2, 4, 8, 16],
24
+ "dec_strides": [2, 4, 5, 6, 8],
25
+ "dec_use_snake": true,
26
+ "dec_final_tanh": false,
27
+ "dec_out_channels": 2,
28
+ "dec_anti_aliasing": false,
29
+ "dec_use_nearest_upsample": false,
30
+ "dec_use_tanh_at_final": false,
31
+ "bottleneck_type": "vae",
32
+ "bottleneck": {"type": "vae"},
33
+ "activation": "snakebeta",
34
+ "snake_logscale": true,
35
+ "anti_aliasing": false,
36
+ "use_cuda_kernel": false,
37
+ "causal": false,
38
+ "padding_mode": "zeros",
39
+ "vocoder_input_dim": 64,
40
+ "latent_mean": null,
41
+ "latent_std": null
42
+ }
sound_tokenizer/diffusion_pytorch_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a4b6da5975bf89f6853fe589ad4752d281ac79fbdfad52ea90537fa080b4b9c2
3
- size 1985176840
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d4c61cde38acfb0cad9048a140c3533750277a8462b19dc08450d9fe1ad9879
3
+ size 1892409600