--- license: apache-2.0 library_name: pytorch pipeline_tag: text-to-image tags: - text-to-image - diffusion - pixel-space - flow-matching - general-vision-learner - vision-foundation-model - depth-estimation - image-restoration - super-resolution - arxiv:2610.09450 ---

Iris-3B

Pixel-space generation & general vision learner.
A 3-billion-parameter model that paints every pixel directly — no VAE, no latent space — and, fine-tuned, estimates depth and restores and upscales images.

Paper Project Page Code Demo License

Generative priors are a promising foundation for downstream vision tasks. In this project we explore pixel-space generative models as an alternative to vision foundation models such as DINOv2. Most image generators work in a compressed "latent" space and rely on a separate decoder to turn that into an image. Iris-3B skips that step: the network itself outputs every pixel, so nothing is lost to a lossy, texture-biased latent. You type a prompt, and you get a 1024-pixel image. Iris-3B is not only an image generator. The same pixel-space prior, fine-tuned with no architectural change, is a general vision learner for dense tasks where detail matters: monocular depth estimation and image restoration and upscaling ship in this repo (see *A general vision learner* below).
## Examples Generated at native aspect ratios of about one megapixel with the default settings (CFG 3, 100 steps).

Studio portrait of a Maasai elder wearing vibrant beaded jewelry, deep red cloth, dark backdrop, Rembrandt lighting, ultra detailed skin texture

Close-up portrait of a young woman with silver glitter freckles and iridescent makeup, soft pastel background, high fashion beauty photography

A glass sculpture of a heart filled with flowers, caustics and reflections, 3D render

Black and white portrait of a fisherman with a thick grey beard and deep wrinkles, piercing eyes, overcast light, fine grain film photograph

A Byzantine-style mosaic of a peacock made of tiny gold and turquoise tiles, shimmering texture

A bronze sculpture of a horse in motion, patina, dramatic museum spotlight

The ancient city of Petra with the Treasury carved into pink sandstone, morning light

A tree with lightbulbs instead of fruits glowing at dusk, surreal concept art

Thousands of sky lanterns rising into the night sky at Yi Peng festival in Chiang Mai

The aurora borealis swirling green and violet over a snowy Lofoten fishing village with red wooden cabins, reflections in a calm fjord, night photograph

Volcanic eruption at night in Iceland, rivers of glowing lava flowing across black fields, plumes of steam lit orange, long exposure
## Get started **1. Install the code** ```bash git clone https://github.com/speridlabs/iris-3b.git cd iris-3b pip install -e . # Python 3.11+, PyTorch 2.7.1+ ``` **2. Download the weights** (about 12 GB for text-to-image; the `depth/` and `upscaler/` folders add about 12 GB each) ```bash hf download speridlabs/iris-3b --local-dir iris-3b --exclude "depth/*" "upscaler/*" ``` **3. Generate an image** ```bash python scripts/sample.py --checkpoint iris-3b \ --prompt "a red fox sleeping in fresh snow, golden hour" ``` The text encoder (Qwen3-VL-4B-Instruct) downloads automatically on the first run. You need an NVIDIA GPU with CUDA. ## Tips | Setting | Default | What it does | |---|---|---| | `--cfg-scale` | 3 | How strictly the image follows the prompt. Higher values follow the prompt more literally but can look harsher. | | `--steps` | 100 | Number of denoising steps. Fewer steps are faster but lose some detail. | | `--seed` | — | Fix it to get the same image again. | | `--negative-prompt` | — | Things you don't want in the image. | | `--txt-file` | — | A text file with one prompt per line, for batches. | Default output is 1024×1024. Write prompts as plain descriptive English sentences. ## A general vision learner: depth, image restoration and upscaling The same backbone, fine-tuned per task: Iris-3B estimates monocular depth and restores and upscales images. Both run the whole model in a single forward pass with the empty prompt, so no text encoder is needed. The weights download on first use. **Monocular depth** — photo → relative depth:

Photo

Iris-3B depth

Photo

Iris-3B depth
**Image restoration and upscaling** — degraded input → restored image:

Input

Iris-3B

Input

Iris-3B
Try both interactively in the [demo](https://huggingface.co/spaces/speridlabs/iris-3b), or from the command line: ```bash python scripts/depth.py photo.jpg --out depth_out # .npy + colorized .png python scripts/upscale.py photo.jpg --out upscaled # _x4.png ``` | Task | Training | Result | |---|---|---| | Depth | Direct regression of relative log depth, 10K steps | AbsRel 0.071, δ1 0.946 (mean of NYUv2, KITTI, ETH3D, ScanNet, DIODE; Marigold V2 protocol, zero-shot) | | Image restoration and upscaling | One-step adversarial restoration (HYPIR recipe), 10K steps, EMA weights | LPIPS 0.292, PSNR 20.68 (DIV2K validation, 4×) | Depth is affine-invariant, not metric: fit a scale and shift in log space to compare with ground truth. Restoration tiles large outputs into overlapping 1024×1024 windows. ## What's in this repo | File | Contents | |---|---| | `model.safetensors` | Text-to-image weights (EMA, float32; run with bfloat16 autocast) | | `config.yaml` | Architecture and sampling settings read by `scripts/sample.py` | | `depth/` | Depth model: `model.safetensors`, `empty_prompt.safetensors`, `config.yaml` | | `upscaler/` | Image restoration and upscaling model: same three files | ## About the model | | | |---|---| | Size | 3B parameters (plus a frozen 4B text encoder) | | Architecture | Diffusion transformer: 8 dual-stream + 16 single-stream blocks, then a 4-block pixel head that turns each 16×16 patch back into pixels | | Text encoder | Qwen3-VL-4B-Instruct (frozen) | | Training | Rectified flow, trained from scratch at 256 → 512 → 1024 px, then supervised fine-tuning at 1024 px (665K steps total) | The full training recipe is in the [code repository](https://github.com/speridlabs/iris-3b). ## Limitations - Like every image generator, Iris-3B can produce inaccurate, biased or unsafe content. Review outputs before using them. - Text rendering inside images and exact object counts are not always reliable. - Images are generated at about one megapixel; for larger prints, use an upscaler. ## License Released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). The text encoder, Qwen3-VL-4B-Instruct, is downloaded from its publisher and stays under its own license (Apache 2.0). ## Citation ```bibtex @article{licai2026iris, title = {Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning}, author = {Li Cai, Hanqiu and Garabito, Chema}, journal = {arXiv preprint arXiv:2610.09450}, year = {2026}, url = {https://arxiv.org/abs/2610.09450} } ``` **Made with ❤️ by [speridlabs.com](https://speridlabs.com)**