Add pipeline tag and links to paper, project, and code
#2
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -1,20 +1,25 @@
|
|
| 1 |
---
|
|
|
|
| 2 |
license: other
|
| 3 |
license_name: nvidia-one-way-noncommercial-license
|
| 4 |
license_link: https://developer.download.nvidia.com/licenses/NVIDIA-OneWay-Noncommercial-License-22Mar2022.pdf
|
| 5 |
-
|
| 6 |
tags:
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
---
|
| 13 |
|
| 14 |
# PixelUMM
|
| 15 |
|
| 16 |
**PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation**
|
| 17 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 18 |
PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint
|
| 19 |
understanding and generation across text, images, and video directly in pixel
|
| 20 |
space.
|
|
@@ -101,4 +106,4 @@ complying with all applicable upstream licenses and terms.
|
|
| 101 |
## References
|
| 102 |
|
| 103 |
- [Qwen3-8B model repository](https://huggingface.co/Qwen/Qwen3-8B)
|
| 104 |
-
- [BAGEL source repository](https://github.com/ByteDance-Seed/Bagel)
|
|
|
|
| 1 |
---
|
| 2 |
+
base_model: Qwen/Qwen3-8B
|
| 3 |
license: other
|
| 4 |
license_name: nvidia-one-way-noncommercial-license
|
| 5 |
license_link: https://developer.download.nvidia.com/licenses/NVIDIA-OneWay-Noncommercial-License-22Mar2022.pdf
|
| 6 |
+
pipeline_tag: any-to-any
|
| 7 |
tags:
|
| 8 |
+
- multimodal
|
| 9 |
+
- image-generation
|
| 10 |
+
- video-generation
|
| 11 |
+
- image-to-text
|
| 12 |
+
- video-text-to-text
|
| 13 |
---
|
| 14 |
|
| 15 |
# PixelUMM
|
| 16 |
|
| 17 |
**PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation**
|
| 18 |
|
| 19 |
+
**Paper:** [PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation](https://huggingface.co/papers/2609.38597)
|
| 20 |
+
**Project:** [PixelUMM Project Page](https://nv-tlabs.github.io/PixelUMM/)
|
| 21 |
+
**Code:** [GitHub](https://github.com/nv-tlabs/PixelUMM)
|
| 22 |
+
|
| 23 |
PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint
|
| 24 |
understanding and generation across text, images, and video directly in pixel
|
| 25 |
space.
|
|
|
|
| 106 |
## References
|
| 107 |
|
| 108 |
- [Qwen3-8B model repository](https://huggingface.co/Qwen/Qwen3-8B)
|
| 109 |
+
- [BAGEL source repository](https://github.com/ByteDance-Seed/Bagel)
|