Robotics
PyTorch
robot-learning
world-action-model
flow-matching
libero
robocoin
galaxea
PoopBear commited on
Commit
369e061
·
verified ·
1 Parent(s): dee892d

Add LIBERO Cosmos Policy Predict2-2B checkpoint

Browse files

Mirror the verified upstream model, configuration, statistics, text embeddings, original model card and license under libero/cosmos-2. Update dataset indexes and document model-specific licenses.

README.md CHANGED
@@ -1,10 +1,15 @@
1
  ---
2
  pretty_name: RIFT Checkpoints
3
- license: apache-2.0
 
 
4
  library_name: pytorch
5
- base_model: Wan-AI/Wan2.2-TI2V-5B
 
 
6
  datasets:
7
  - yuanty/LIBERO-fastwam
 
8
  tags:
9
  - robotics
10
  - robot-learning
@@ -27,16 +32,22 @@ Each dataset has one top-level directory. Model variants live inside that direct
27
 
28
  | Dataset | Directory | Models |
29
  | --- | --- | --- |
30
- | LIBERO | [libero](https://huggingface.co/PoopBear/RIFT/tree/main/libero) | FastWAM; RIFT, FastWAM-IDM and FastWAM-Joint (step 21,700) |
31
  | RoboCOIN, three tasks | [robocoin_multitask_mm](https://huggingface.co/PoopBear/RIFT/tree/main/robocoin_multitask_mm) | RIFT (final), 10 epochs / step 9,340 |
32
  | Galaxea indoor cleaning | [galaxea_indoor_cleaning_3cam224](https://huggingface.co/PoopBear/RIFT/tree/main/galaxea_indoor_cleaning_3cam224) | FastWAM, FastWAM-Joint and RIFT, 10 epochs / step 14,830 |
33
 
34
- Shared VAE, text-encoder, and tokenizer files are under `assets/`. Exact checkpoint paths,
35
  hashes, and source revisions are listed in [checkpoint_manifest.json](checkpoint_manifest.json).
36
  Use each model's own normalization statistics.
37
 
38
  ## Download
39
 
 
 
 
 
 
 
40
  LIBERO FastWAM:
41
 
42
  ```bash
@@ -79,5 +90,6 @@ with `weights_only=True`.
79
 
80
  ## License
81
 
82
- The checkpoint is distributed under the Apache License 2.0. RIFT source code
83
- is distributed separately under the MIT License. See `NOTICE` for attribution.
 
 
1
  ---
2
  pretty_name: RIFT Checkpoints
3
+ license: other
4
+ license_name: model-specific-licenses
5
+ license_link: https://huggingface.co/PoopBear/RIFT/blob/main/README.md#license
6
  library_name: pytorch
7
+ base_model:
8
+ - Wan-AI/Wan2.2-TI2V-5B
9
+ - nvidia/Cosmos-Predict2-2B-Video2World
10
  datasets:
11
  - yuanty/LIBERO-fastwam
12
+ - nvidia/LIBERO-Cosmos-Policy
13
  tags:
14
  - robotics
15
  - robot-learning
 
32
 
33
  | Dataset | Directory | Models |
34
  | --- | --- | --- |
35
+ | LIBERO | [libero](https://huggingface.co/PoopBear/RIFT/tree/main/libero) | FastWAM; RIFT, FastWAM-IDM and FastWAM-Joint (step 21,700); Cosmos Policy Predict2-2B |
36
  | RoboCOIN, three tasks | [robocoin_multitask_mm](https://huggingface.co/PoopBear/RIFT/tree/main/robocoin_multitask_mm) | RIFT (final), 10 epochs / step 9,340 |
37
  | Galaxea indoor cleaning | [galaxea_indoor_cleaning_3cam224](https://huggingface.co/PoopBear/RIFT/tree/main/galaxea_indoor_cleaning_3cam224) | FastWAM, FastWAM-Joint and RIFT, 10 epochs / step 14,830 |
38
 
39
+ Shared Wan VAE, text-encoder, and tokenizer files are under `assets/`. Exact checkpoint paths,
40
  hashes, and source revisions are listed in [checkpoint_manifest.json](checkpoint_manifest.json).
41
  Use each model's own normalization statistics.
42
 
43
  ## Download
44
 
45
+ LIBERO Cosmos-2:
46
+
47
+ ```bash
48
+ hf download PoopBear/RIFT --include 'libero/cosmos-2/*' --local-dir ./checkpoints
49
+ ```
50
+
51
  LIBERO FastWAM:
52
 
53
  ```bash
 
90
 
91
  ## License
92
 
93
+ RIFT checkpoints are distributed under the Apache License 2.0 (see `LICENSE`). RIFT source code is distributed separately under the MIT License. See `NOTICE` for attribution.
94
+
95
+ The third-party Cosmos Policy files in `libero/cosmos-2/` retain their original NVIDIA One-Way Noncommercial License (NSCLv1). See [the included license](libero/cosmos-2/LICENSE) and [original model card](libero/cosmos-2/UPSTREAM_README.md).
checkpoint_manifest.json CHANGED
@@ -137,6 +137,36 @@
137
  "checkpoint": "libero_uncond_2cam224.pt",
138
  "dataset_stats": "libero_uncond_2cam224_dataset_stats.json"
139
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
140
  }
141
  ]
142
  }
 
137
  "checkpoint": "libero_uncond_2cam224.pt",
138
  "dataset_stats": "libero_uncond_2cam224_dataset_stats.json"
139
  }
140
+ },
141
+ {
142
+ "dataset": "libero",
143
+ "variant": "cosmos-2",
144
+ "model_name": "Cosmos-Policy-LIBERO-Predict2-2B",
145
+ "checkpoint": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
146
+ "bytes": 3913017345,
147
+ "sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2",
148
+ "step": 40000,
149
+ "step_source": "upstream config.json training.gradient_steps",
150
+ "config": "libero/cosmos-2/config.json",
151
+ "config_kind": "original upstream model metadata",
152
+ "dataset_stats": "libero/cosmos-2/libero_dataset_statistics.json",
153
+ "text_embeddings": "libero/cosmos-2/libero_t5_embeddings.pkl",
154
+ "source": {
155
+ "repo": "nvidia/Cosmos-Policy-LIBERO-Predict2-2B",
156
+ "revision": "cb689ec0e3347c13667d70a78a3447388f5c3bb8",
157
+ "path": "Cosmos-Policy-LIBERO-Predict2-2B.pt"
158
+ },
159
+ "license": {
160
+ "name": "NVIDIA One-Way Noncommercial License (NSCLv1)",
161
+ "path": "libero/cosmos-2/LICENSE",
162
+ "source": {
163
+ "repo": "NVlabs/HMAR",
164
+ "revision": "7e17e31191ee0287c3043e3a8523c23ee3d355dd",
165
+ "url": "https://github.com/NVlabs/HMAR/blob/7e17e31191ee0287c3043e3a8523c23ee3d355dd/LICENSE",
166
+ "bytes": 4060,
167
+ "sha256": "6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5"
168
+ }
169
+ }
170
  }
171
  ]
172
  }
libero/README.md CHANGED
@@ -2,9 +2,12 @@
2
 
3
  | Model | Directory | Step |
4
  | --- | --- | --- |
 
5
  | FastWAM | [fastwam](fastwam/) | Not recorded |
6
  | RIFT | [rift](rift/) | 21,700 |
7
  | FastWAM-IDM | [idm](idm/) | 21,700 |
8
  | FastWAM-Joint | [joint](joint/) | 21,700 |
9
 
10
- Each directory contains the original weights, matching normalization statistics, reference configs, provenance, and SHA-256 checksums. RIFT's configuration is a model configuration; the FastWAM, IDM and Joint configs are workspace references.
 
 
 
2
 
3
  | Model | Directory | Step |
4
  | --- | --- | --- |
5
+ | Cosmos Policy Predict2-2B | [cosmos-2](cosmos-2/) | 40,000 (upstream config) |
6
  | FastWAM | [fastwam](fastwam/) | Not recorded |
7
  | RIFT | [rift](rift/) | 21,700 |
8
  | FastWAM-IDM | [idm](idm/) | 21,700 |
9
  | FastWAM-Joint | [joint](joint/) | 21,700 |
10
 
11
+ Each directory contains the original weights, matching normalization statistics, configuration files, provenance, and SHA-256 checksums. RIFT's configuration is a model configuration; the FastWAM, IDM and Joint configs are workspace references.
12
+
13
+ Cosmos-2 includes the original upstream model metadata, cached text embeddings, model card, and NVIDIA license.
libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2
3
+ size 3913017345
libero/cosmos-2/LICENSE ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ NVIDIA License
2
+
3
+ 1. Definitions
4
+
5
+ “Licensor” means any person or entity that distributes its Work.
6
+ “Work” means (a) the original work of authorship made available under this license, which may include software, documentation, or other files, and (b) any additions to or derivative works thereof that are made available under this license.
7
+ The terms “reproduce,” “reproduction,” “derivative works,” and “distribution” have the meaning as provided under U.S. copyright law; provided, however, that for the purposes of this license, derivative works shall not include works that remain separable from, or merely link (or bind by name) to the interfaces of, the Work.
8
+ Works are “made available” under this license by including in or with the Work either (a) a copyright notice referencing the applicability of this license to the Work, or (b) a copy of this license.
9
+
10
+ 2. License Grant
11
+
12
+ 2.1 Copyright Grant. Subject to the terms and conditions of this license, each Licensor grants to you a perpetual, worldwide, non-exclusive, royalty-free, copyright license to use, reproduce, prepare derivative works of, publicly display, publicly perform, sublicense and distribute its Work and any resulting derivative works in any form.
13
+
14
+ 3. Limitations
15
+
16
+ 3.1 Redistribution. You may reproduce or distribute the Work only if (a) you do so under this license, (b) you include a complete copy of this license with your distribution, and (c) you retain without modification any copyright, patent, trademark, or attribution notices that are present in the Work.
17
+
18
+ 3.2 Derivative Works. You may specify that additional or different terms apply to the use, reproduction, and distribution of your derivative works of the Work (“Your Terms”) only if (a) Your Terms provide that the use limitation in Section 3.3 applies to your derivative works, and (b) you identify the specific derivative works that are subject to Your Terms. Notwithstanding Your Terms, this license (including the redistribution requirements in Section 3.1) will continue to apply to the Work itself.
19
+
20
+ 3.3 Use Limitation. The Work and any derivative works thereof only may be used or intended for use non-commercially. Notwithstanding the foregoing, NVIDIA Corporation and its affiliates may use the Work and any derivative works commercially. As used herein, “non-commercially” means for non-commercial research and educational purposes only.
21
+
22
+ 3.4 Patent Claims. If you bring or threaten to bring a patent claim against any Licensor (including any claim, cross-claim or counterclaim in a lawsuit) to enforce any patents that you allege are infringed by any Work, then your rights under this license from such Licensor (including the grant in Section 2.1) will terminate immediately.
23
+
24
+ 3.5 Trademarks. This license does not grant any rights to use any Licensor’s or its affiliates’ names, logos, or trademarks, except as necessary to reproduce the notices described in this license.
25
+
26
+ 3.6 Termination. If you violate any term of this license, then your rights under this license (including the grant in Section 2.1) will terminate immediately.
27
+
28
+ 4. Disclaimer of Warranty.
29
+
30
+ THE WORK IS PROVIDED “AS IS” WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING WARRANTIES OR CONDITIONS OF
31
+ MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE OR NON-INFRINGEMENT. YOU BEAR THE RISK OF UNDERTAKING ANY ACTIVITIES UNDER THIS LICENSE.
32
+
33
+ 5. Limitation of Liability.
34
+
35
+ EXCEPT AS PROHIBITED BY APPLICABLE LAW, IN NO EVENT AND UNDER NO LEGAL THEORY, WHETHER IN TORT (INCLUDING NEGLIGENCE), CONTRACT, OR OTHERWISE SHALL ANY LICENSOR BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES ARISING OUT OF OR RELATED TO THIS LICENSE, THE USE OR INABILITY TO USE THE WORK (INCLUDING BUT NOT LIMITED TO LOSS OF GOODWILL, BUSINESS INTERRUPTION, LOST PROFITS OR DATA, COMPUTER FAILURE OR MALFUNCTION, OR ANY OTHER DAMAGES OR LOSSES), EVEN IF THE LICENSOR HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
libero/cosmos-2/README.md ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # LIBERO Cosmos-2
2
+
3
+ This directory mirrors **Cosmos-Policy-LIBERO-Predict2-2B**, the NVIDIA Cosmos Policy Predict2-2B checkpoint for LIBERO. All original model files match both the verified server bundle and the pinned upstream revision.
4
+
5
+ - Weights: `Cosmos-Policy-LIBERO-Predict2-2B.pt`.
6
+ - Original model metadata: `config.json`.
7
+ - Normalization: `libero_dataset_statistics.json`.
8
+ - Cached text embeddings: `libero_t5_embeddings.pkl`.
9
+ - Original model card: [UPSTREAM_README.md](UPSTREAM_README.md).
10
+ - File hashes and source revisions: `SHA256SUMS` and `provenance.json`.
11
+
12
+ Upstream: [nvidia/Cosmos-Policy-LIBERO-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B/tree/cb689ec0e3347c13667d70a78a3447388f5c3bb8). Use the upstream Cosmos Policy code and its required pretrained assets for inference.
13
+
14
+ ```bash
15
+ hf download PoopBear/RIFT --include 'libero/cosmos-2/*' --local-dir ./checkpoints
16
+ ```
17
+
18
+ ## License
19
+
20
+ These Cosmos files retain the NVIDIA One-Way Noncommercial License (NSCLv1), as specified by the upstream model card. A complete copy is included in [LICENSE](LICENSE). The original model card and attribution notices are preserved.
libero/cosmos-2/SHA256SUMS ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ 8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2 Cosmos-Policy-LIBERO-Predict2-2B.pt
2
+ 3bd882474403177ab22e65598d8ab1c1b299804b5bff0019f30c4a01aa405506 UPSTREAM_README.md
3
+ 6c246e6588762546fd91c4fac62af570583da1156aa9b2800f2268e6a2c31f43 config.json
4
+ 5b119a98ad7824507ddff3c6c7507ca244261ee416dcfa623876702160c580d3 libero_dataset_statistics.json
5
+ 8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5 libero_t5_embeddings.pkl
6
+ 6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5 LICENSE
7
+ ab9b5f6d5f126b63bcf5f20c24594f36ae2bb490ffaf625845af8f689adb9a5c README.md
8
+ bd742a925ea01fda83403adc90aa915b96fdfed700dd374b8212a455194da70d provenance.json
libero/cosmos-2/UPSTREAM_README.md ADDED
@@ -0,0 +1,236 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model:
3
+ - nvidia/Cosmos-Predict2-2B-Video2World
4
+ ---
5
+ # **Cosmos-Policy-LIBERO-Predict2-2B**
6
+
7
+ [**Cosmos Policy**](https://huggingface.co/collections/nvidia/cosmos-policy) | [**Code**](http://github.com/NVlabs/cosmos-policy) | [**White Paper**](https://arxiv.org/abs/2601.16163) | [**Website**](https://research.nvidia.com/labs/dir/cosmos-policy/)
8
+
9
+ # Model Overview
10
+
11
+ ## Description:
12
+
13
+ Cosmos-Policy-LIBERO-Predict2-2B is a 2B-parameter robot manipulation policy model fine-tuned from the [NVIDIA Cosmos-Predict2-2B-Video2World](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World) video foundation model. This model achieves state-of-the-art performance on the LIBERO simulation benchmark with a 98.5% average success rate across four task suites.
14
+
15
+ Key features:
16
+
17
+ * **Single-stage fine-tuning**: Adapted from pretrained video model with no architectural modifications
18
+ * **Multimodal outputs**: Jointly predicts actions, future states, and values through unified video diffusion
19
+ * **High performance**: 98.5% average success rate on LIBERO (Spatial: 98.1%, Object: 100.0%, Goal: 98.2%, Long: 97.6%)
20
+
21
+ Use cases:
22
+
23
+ * Robotic manipulation and control in simulation environments
24
+ * Imitation learning and policy learning for table-top manipulation tasks
25
+ * Vision-based robot learning with multiple camera viewpoints
26
+ * Long-horizon task planning and execution
27
+ * Lifelong learning and transfer learning in robotics
28
+
29
+ This model is for research and development only.
30
+
31
+ **Model Developer**: NVIDIA
32
+
33
+ ## Model Versions
34
+
35
+ Cosmos Policy models include the following:
36
+
37
+ - [Cosmos-Policy-LIBERO-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in simulated LIBERO environments.
38
+ - [Cosmos-Policy-RoboCasa-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-RoboCasa-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in simulated RoboCasa environments.
39
+ - [Cosmos-Policy-ALOHA-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-ALOHA-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in real-world ALOHA robot environments.
40
+ - [Cosmos-Policy-ALOHA-Planning-Model-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-ALOHA-Planning-Model-Predict2-2B): Given current state observations, a task description, and action sequences, generate future state predictions and value estimates for robot manipulation in real-world ALOHA robot environments. (This checkpoint is meant to be deployed alongside Cosmos-Policy-ALOHA-Predict2-2B, not independently.)
41
+
42
+ ### License:
43
+
44
+ This model is released under the [NVIDIA One-Way Noncommercial License (NSCLv1)](https://github.com/NVlabs/HMAR/blob/main/LICENSE). For a custom license, please contact [cosmos-license@nvidia.com](mailto:cosmos-license@nvidia.com).
45
+
46
+ Under the NVIDIA One-Way Noncommercial License (NSCLv1), NVIDIA confirms:
47
+
48
+ * Models are not for commercial use.
49
+ * NVIDIA does not claim ownership to any outputs generated using the Models or Derivative Models.
50
+
51
+ ### Deployment Geography:
52
+
53
+ Global
54
+
55
+ ### Use Case:
56
+
57
+ Physical AI: Robot manipulation and control, encompassing tabletop manipulation and imitation learning in simulation environments.
58
+
59
+ ### Release Date:
60
+
61
+ GitHub [01/22/2026] via [https://github.com/nvlabs/cosmos-policy](https://github.com/nvlabs/cosmos-policy)
62
+
63
+ Hugging Face [01/22/2026] via [https://huggingface.co/collections/nvidia/cosmos-policy](https://huggingface.co/collections/nvidia/cosmos-policy)
64
+
65
+ ## Model Architecture:
66
+
67
+ Architecture Type: A diffusion transformer with latent video diffusion, fine-tuned from Cosmos-Predict2-2B-Video2World.
68
+
69
+ Network Architecture: The model uses the same architecture as the base [Cosmos-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World) model (a diffusion transformer with latent video diffusion).
70
+
71
+ **Key adaptation**: Actions, proprioceptive states, and values are encoded as latent frames and injected directly into the video model's latent diffusion sequence, enabling the model to generate these modalities alongside predicted future images.
72
+
73
+ **Number of model parameters:**
74
+
75
+ 2B (inherited from base model)
76
+
77
+ ## Input
78
+
79
+ **Input Type(s)**: Text + Multi-view Images + Proprioceptive State
80
+
81
+ **Input Format(s)**:
82
+
83
+ * Text: String (natural language task description)
84
+ * Images: RGB images from multiple camera views
85
+ * Proprioception: Numerical array
86
+
87
+ **Input Parameters**:
88
+
89
+ * Text: One-dimensional (1D) - Task description (e.g., "put the black bowl on top of the cabinet")
90
+ * Images: Two-dimensional (2D) - Third-person camera (agentview): 224×224 RGB; Wrist-mounted camera (eye-in-hand): 224×224 RGB
91
+ * Proprioception: One-dimensional (1D) - 9-dimensional state (2 gripper joints + 3 end-effector position + 4 end-effector quaternion)
92
+
93
+ **Other Properties Related to Input**:
94
+
95
+ * Requires specific camera configuration (third-person + wrist views)
96
+ * Images resized to 224×224 pixels from original resolution
97
+ * Trained exclusively for Franka Emika Panda robot arm in LIBERO simulation environments
98
+
99
+ ## Output
100
+
101
+ **Output Type(s)**: Action Sequence + Future State Predictions + Value Estimate
102
+
103
+ **Output Format**:
104
+
105
+ * Actions: Numerical array
106
+ * Future states: Images + Proprioception
107
+ * Value: Scalar
108
+
109
+ **Output Parameters**:
110
+
111
+ * Action chunk: 16-timestep sequence of 7-dimensional actions (6-DoF end-effector control + 1 gripper)
112
+ * Future robot proprioception: 9-dimensional state at timestep t+16
113
+ * Future state images: Third-person camera prediction (224×224 RGB) and wrist camera prediction (224×224 RGB) at timestep t+16
114
+ * Future state value: Expected cumulative reward from future state (scalar)
115
+
116
+ **Other Properties Related to Output**:
117
+
118
+ * Action chunk size: 16 timesteps
119
+ * Denoising steps: 5 (configurable without retraining)
120
+ * Noise level range: σ_min = 4.0, σ_max = 80.0
121
+ * Generation mode: Parallel (action, future state, and value generated simultaneously)
122
+
123
+ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
124
+
125
+ ## Software Integration
126
+
127
+ **Runtime Engine(s):**
128
+
129
+ * [Transformers](https://github.com/huggingface/transformers)
130
+
131
+ **Supported Hardware Microarchitecture Compatibility:**
132
+
133
+ * NVIDIA Hopper (e.g., H100)
134
+
135
+ **Note**: We have only tested doing inference with BF16 precision.
136
+
137
+ **Operating System(s):**
138
+
139
+ * Linux
140
+
141
+ The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
142
+
143
+ # Usage
144
+
145
+ See [Cosmos Policy GitHub](http://github.com/NVlabs/cosmos-policy) for details.
146
+
147
+ ## Training and Evaluation Sections:
148
+
149
+ ### Training Datasets:
150
+
151
+ **Data Collection Method**:
152
+
153
+ * LIBERO-Cosmos-Policy: Hybrid: Human - Human-teleoperated demonstrations recorded in simulation environment
154
+
155
+ **Labeling Method**:
156
+
157
+ * LIBERO-Cosmos-Policy: Automated - Success/failure labels automatically determined by simulation environment evaluation; task descriptions from benchmark specification
158
+
159
+ ##### Properties:
160
+
161
+ **Training Data**: [LIBERO-Cosmos-Policy](https://huggingface.co/datasets/nvidia/LIBERO-Cosmos-Policy) dataset
162
+
163
+ - 4 task suites: LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, LIBERO-Long
164
+ - 500 demonstrations per suite (50 demos × 10 tasks)
165
+ - Successful demonstrations used for policy training
166
+ - All demonstrations (including failures) used for world model and value function training
167
+
168
+ **Training Configuration**:
169
+
170
+ - **Base model**: NVIDIA Cosmos-Predict2-2B-Video2World (`model-480p-16fps.pt`)
171
+ - **Training steps**: 40,000 gradient steps
172
+ - **Batch size**: 1,920 (global)
173
+ - **GPUs**: 64 H100 GPUs
174
+ - **Training time**: ~48 hours
175
+ - **Optimization**: Full model fine-tuning (all weights updated)
176
+ - **Action chunk size**: 16 timesteps
177
+ - **Image resolution**: 224×224 pixels
178
+
179
+ **Training Objective**: The model is trained with a hybrid log-normal-uniform noise distribution (modified from the base model's log-normal distribution; see paper for details) to improve action prediction accuracy. Training batches are split 50/25/25 for policy, world model, and value function objectives, respectively.
180
+
181
+ ### Evaluation Datasets:
182
+
183
+ Data Collection Method: Not Applicable
184
+
185
+ Labeling Method: Not Applicable
186
+
187
+ Properties: Not Applicable - we use the LIBERO simulation environments for direct evaluations.
188
+
189
+ ## Inference:
190
+
191
+ **Test Hardware:** H100, A100
192
+
193
+ See [Cosmos Policy GitHub](http://github.com/NVlabs/cosmos-policy) for details.
194
+
195
+ #### System Requirements and Performance
196
+
197
+ Inference with base Cosmos Policy only (i.e., no model-based planning):
198
+
199
+ * 1 GPU with 6.8 GB VRAM for LIBERO sim benchmark tasks
200
+ * 1 GPU with 8.9 GB VRAM for RoboCasa sim benchmark tasks
201
+ * 1 GPU with 6.0 GB VRAM for ALOHA robot tasks
202
+
203
+ #### Quality Benchmarks
204
+
205
+ ### LIBERO Benchmark Results
206
+
207
+ | Task Suite | Success Rate |
208
+ | ----------------- | --------------- |
209
+ | LIBERO-Spatial | 98.1% |
210
+ | LIBERO-Object | 100.0% |
211
+ | LIBERO-Goal | 98.2% |
212
+ | LIBERO-Long | 97.6% |
213
+ | **Average** | **98.5%** |
214
+
215
+ Success rates are averaged over 500 trials per suite (10 tasks × 50 episodes) across 3 random seeds (6,000 trials total).
216
+
217
+ ## Ethical Considerations
218
+
219
+ NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
220
+
221
+ Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.
222
+
223
+ Please report security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
224
+
225
+ ## Related Resources
226
+
227
+ - **Base Model**: [Cosmos-Predict2-2B-Video2World](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World)
228
+ - **Training Dataset**: [LIBERO-Cosmos-Policy](https://huggingface.co/datasets/nvidia/LIBERO-Cosmos-Policy)
229
+ - **Paper**: [Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning](https://arxiv.org/abs/2601.16163)
230
+ - **Original LIBERO**: [LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning](https://arxiv.org/abs/2306.03310)
231
+
232
+ ## Citation
233
+
234
+ If you use this model, please cite the Cosmos Policy paper:
235
+
236
+ (Cosmos Policy BibTeX citation coming soon!)
libero/cosmos-2/config.json ADDED
@@ -0,0 +1,67 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "cosmos-policy",
3
+ "architecture": "diffusion-transformer",
4
+ "base_model": "nvidia/Cosmos-Predict2-2B-Video2World",
5
+ "num_parameters": "2B",
6
+
7
+ "input_spec": {
8
+ "text": {
9
+ "type": "string",
10
+ "description": "Natural language task description"
11
+ },
12
+ "images": {
13
+ "format": "RGB",
14
+ "resolution": [224, 224],
15
+ "views": ["agentview", "eye_in_hand"]
16
+ },
17
+ "proprioception": {
18
+ "dim": 9,
19
+ "components": ["gripper_joints", "end_effector_position", "quaternion"]
20
+ }
21
+ },
22
+
23
+ "output_spec": {
24
+ "actions": {
25
+ "dim": 7,
26
+ "horizon": 16,
27
+ "components": ["end_effector_6dof", "gripper"]
28
+ },
29
+ "future_proprioception": {
30
+ "dim": 9
31
+ },
32
+ "future_images": {
33
+ "resolution": [224, 224]
34
+ },
35
+ "value": {
36
+ "dim": 1
37
+ }
38
+ },
39
+
40
+ "diffusion_config": {
41
+ "denoising_steps": 5,
42
+ "sigma_min": 4.0,
43
+ "sigma_max": 80.0,
44
+ "generation_mode": "parallel"
45
+ },
46
+
47
+ "training": {
48
+ "dataset": "LIBERO-Cosmos-Policy",
49
+ "gradient_steps": 40000,
50
+ "batch_size": 1920,
51
+ "hardware": "64x H100",
52
+ "action_chunk_size": 16
53
+ },
54
+
55
+ "benchmark_results": {
56
+ "libero_spatial": 0.981,
57
+ "libero_object": 1.0,
58
+ "libero_goal": 0.982,
59
+ "libero_long": 0.976,
60
+ "average": 0.985
61
+ },
62
+
63
+ "inference": {
64
+ "precision": "bf16",
65
+ "vram_gb": 6.8
66
+ }
67
+ }
libero/cosmos-2/libero_dataset_statistics.json ADDED
@@ -0,0 +1,102 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "actions_min": [
3
+ -0.9375,
4
+ -0.9375,
5
+ -0.9375,
6
+ -0.2582142949104309,
7
+ -0.375,
8
+ -0.3642857074737549,
9
+ -1.0
10
+ ],
11
+ "actions_max": [
12
+ 0.9375,
13
+ 0.9375,
14
+ 0.9375,
15
+ 0.3557142913341522,
16
+ 0.375,
17
+ 0.375,
18
+ 1.0
19
+ ],
20
+ "actions_mean": [
21
+ 0.0627586841583252,
22
+ 0.08706478029489517,
23
+ -0.09038306772708893,
24
+ 0.0004062088264618069,
25
+ 0.005638255272060633,
26
+ -0.004925117362290621,
27
+ -0.0539495050907135
28
+ ],
29
+ "actions_std": [
30
+ 0.3358899652957916,
31
+ 0.3785707950592041,
32
+ 0.44441089034080505,
33
+ 0.03959346562623978,
34
+ 0.06336744129657745,
35
+ 0.07815399020910263,
36
+ 0.9980894923210144
37
+ ],
38
+ "actions_median": [
39
+ 0.0,
40
+ 0.0,
41
+ -0.05624999850988388,
42
+ 0.0,
43
+ 0.0,
44
+ 0.0,
45
+ -1.0
46
+ ],
47
+ "proprio_min": [
48
+ -0.005054043605923653,
49
+ -0.042120561003685,
50
+ -0.48564884066581726,
51
+ -0.33136284351348877,
52
+ 0.008128181099891663,
53
+ 0.2902947962284088,
54
+ -0.8897988200187683,
55
+ -0.5616180300712585,
56
+ -0.5712438225746155
57
+ ],
58
+ "proprio_max": [
59
+ 0.04238177835941315,
60
+ 0.0013513736193999648,
61
+ 0.2103137969970703,
62
+ 0.3904264271259308,
63
+ 1.3660907745361328,
64
+ 0.9999998807907104,
65
+ 0.9094555377960205,
66
+ 0.34964922070503235,
67
+ 0.6364591121673584
68
+ ],
69
+ "proprio_mean": [
70
+ 0.026981549337506294,
71
+ -0.027262669056653976,
72
+ -0.04627704992890358,
73
+ 0.03411973640322685,
74
+ 0.7628858089447021,
75
+ 0.9422642588615417,
76
+ -0.05906042829155922,
77
+ -0.04012482985854149,
78
+ -0.0029695192351937294
79
+ ],
80
+ "proprio_std": [
81
+ 0.014144331216812134,
82
+ 0.014038173481822014,
83
+ 0.10425098240375519,
84
+ 0.15156729519367218,
85
+ 0.37788110971450806,
86
+ 0.11462057381868362,
87
+ 0.2741202116012573,
88
+ 0.10206353664398193,
89
+ 0.0958578810095787
90
+ ],
91
+ "proprio_median": [
92
+ 0.03365287557244301,
93
+ -0.03424428775906563,
94
+ -0.02911928854882717,
95
+ 0.02497534640133381,
96
+ 0.9323337078094482,
97
+ 0.9929837584495544,
98
+ -0.023691711947321892,
99
+ -0.027851801365613937,
100
+ 0.0008408676367253065
101
+ ]
102
+ }
libero/cosmos-2/libero_t5_embeddings.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5
3
+ size 41957938
libero/cosmos-2/provenance.json ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "dataset": "libero",
3
+ "variant": "cosmos-2",
4
+ "model_name": "Cosmos-Policy-LIBERO-Predict2-2B",
5
+ "checkpoint": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
6
+ "bytes": 3913017345,
7
+ "sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2",
8
+ "step": 40000,
9
+ "step_source": "upstream config.json training.gradient_steps",
10
+ "config": "libero/cosmos-2/config.json",
11
+ "config_kind": "original upstream model metadata",
12
+ "dataset_stats": "libero/cosmos-2/libero_dataset_statistics.json",
13
+ "text_embeddings": "libero/cosmos-2/libero_t5_embeddings.pkl",
14
+ "source": {
15
+ "repo": "nvidia/Cosmos-Policy-LIBERO-Predict2-2B",
16
+ "revision": "cb689ec0e3347c13667d70a78a3447388f5c3bb8",
17
+ "path": "Cosmos-Policy-LIBERO-Predict2-2B.pt"
18
+ },
19
+ "license": {
20
+ "name": "NVIDIA One-Way Noncommercial License (NSCLv1)",
21
+ "path": "libero/cosmos-2/LICENSE",
22
+ "source": {
23
+ "repo": "NVlabs/HMAR",
24
+ "revision": "7e17e31191ee0287c3043e3a8523c23ee3d355dd",
25
+ "url": "https://github.com/NVlabs/HMAR/blob/7e17e31191ee0287c3043e3a8523c23ee3d355dd/LICENSE",
26
+ "bytes": 4060,
27
+ "sha256": "6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5"
28
+ }
29
+ },
30
+ "files": [
31
+ {
32
+ "source_name": "Cosmos-Policy-LIBERO-Predict2-2B.pt",
33
+ "path": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
34
+ "bytes": 3913017345,
35
+ "sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2"
36
+ },
37
+ {
38
+ "source_name": "README.md",
39
+ "path": "libero/cosmos-2/UPSTREAM_README.md",
40
+ "bytes": 11233,
41
+ "sha256": "3bd882474403177ab22e65598d8ab1c1b299804b5bff0019f30c4a01aa405506"
42
+ },
43
+ {
44
+ "source_name": "config.json",
45
+ "path": "libero/cosmos-2/config.json",
46
+ "bytes": 1354,
47
+ "sha256": "6c246e6588762546fd91c4fac62af570583da1156aa9b2800f2268e6a2c31f43"
48
+ },
49
+ {
50
+ "source_name": "libero_dataset_statistics.json",
51
+ "path": "libero/cosmos-2/libero_dataset_statistics.json",
52
+ "bytes": 2372,
53
+ "sha256": "5b119a98ad7824507ddff3c6c7507ca244261ee416dcfa623876702160c580d3"
54
+ },
55
+ {
56
+ "source_name": "libero_t5_embeddings.pkl",
57
+ "path": "libero/cosmos-2/libero_t5_embeddings.pkl",
58
+ "bytes": 41957938,
59
+ "sha256": "8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5"
60
+ }
61
+ ]
62
+ }