PixelUMM / README.md
bryanchunvidia's picture
Recreate PixelUMM preserving all model files without demo artifacts
b621a91
|
Raw History Blame Contribute Delete
3.88 kB
---
license: other
license_name: nvidia-one-way-noncommercial-license
license_link: https://developer.download.nvidia.com/licenses/NVIDIA-OneWay-Noncommercial-License-22Mar2022.pdf
base_model: Qwen/Qwen3-8B
tags:
- multimodal
- image-generation
- video-generation
- image-to-text
- video-text-to-text
---
# PixelUMM
**PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation**
PixelUMM is an NVIDIA-developed, encoder-free unified multimodal model for joint
understanding and generation across text, images, and video directly in pixel
space.
![PixelUMM architecture](assets/pixelumm-teaser.png)
> PixelUMM is released for non-commercial research or evaluation purposes only.
## Overview
Images are represented as pixel patches instead of embeddings from a separate pretrained
vision encoder, allowing understanding and generation to share a Transformer-based
representation.
PixelUMM supports:
- text-to-image generation;
- image-conditioned text generation;
- video-conditioned understanding; and
- video generation.
## Model Architecture
- **Architecture:** Decoder-only Transformer with raw-pixel patch embeddings and
an iterative pixel-generation head.
- **Language backbone:**
[Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B), revision
`b968826d9c46dd6066d109eabc6255188de91218`.
- **Image representation:** RGB images divided into 16-by-16 pixel patches.
- **Generation:** Iterative denoising for image and video generation.
- **Parameters:** 15,199,672,064.
## Inputs and Outputs
PixelUMM accepts text, RGB images, and video frames. It produces text, RGB images,
and video frames.
## Installation
PixelUMM requires Linux, an NVIDIA GPU, a CUDA build of PyTorch, and
FlashAttention. Install PyTorch and FlashAttention for your CUDA environment, then
install the remaining dependencies.
## Intended Use
PixelUMM is intended for researchers and developers studying unified multimodal
modeling, pixel-space representation learning, multimodal understanding, and image
or video generation. Example uses include:
- studying shared representations across text, images, and video;
- evaluating encoder-free multimodal architectures;
- generating images or video from text prompts; and
- producing text conditioned on images or video.
PixelUMM has not been validated for production or high-stakes use.
## Limitations and Safety
PixelUMM may produce inaccurate, offensive, or otherwise inappropriate content.
It may not follow prompts reliably and may produce semantically or temporally
inconsistent outputs. Performance may vary across languages, visual domains,
video duration, resolution, aspect ratio, and hardware.
Users should implement appropriate safety measures, including content filtering,
abuse monitoring, and access controls. Users are responsible for model inputs and
outputs, for obtaining the rights and permissions required for input content, and
for complying with applicable laws and regulations.
## License
The PixelUMM source code is licensed under the
[Apache License 2.0](LICENSE). Third-party notices are provided in
[THIRD_PARTY_LICENSES.md](THIRD_PARTY_LICENSES.md).
The PixelUMM model checkpoint is a separate artifact licensed under the
[NVIDIA One-Way Noncommercial License](https://developer.download.nvidia.com/licenses/NVIDIA-OneWay-Noncommercial-License-22Mar2022.pdf).
Its use is limited to non-commercial research or evaluation purposes.
PixelUMM uses the [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) base model,
which is licensed under Apache-2.0, and includes modified code derived from
[BAGEL](https://github.com/ByteDance-Seed/Bagel). Users are responsible for
complying with all applicable upstream licenses and terms.
## References
- [Qwen3-8B model repository](https://huggingface.co/Qwen/Qwen3-8B)
- [BAGEL source repository](https://github.com/ByteDance-Seed/Bagel)