Iris-3B for ComfyUI
ComfyUI-ready files for Iris-3B by Speridlabs, a 3B text-to-image transformer that works directly in pixel space (no VAE).
Use them with ComfyUI-Iris. It adds native Iris support without any custom nodes and has an example workflow for each file below.
Files
Text to image
| File | Size | Notes |
|---|---|---|
iris-3b_bf16.safetensors |
6.0 GB | Recommended. |
iris-3b_int8_convrot.safetensors |
3.4 GB | About 1.7x faster. Needs an RTX 20-series or newer. |
iris-3b.safetensors |
12 GB | The original FP32 release, unchanged. |
Depth
| File | Size | Notes |
|---|---|---|
iris-3b-depth_bf16.safetensors |
6.0 GB | Iris-3B's monocular depth fine-tune. |
iris-3b-depth_int8_convrot.safetensors |
3.4 GB | int8 version. |
Upscaling
| File | Size | Notes |
|---|---|---|
loras/iris-3b-upscaler_r128.safetensors |
0.5 GB | Iris-3B's restoration fine-tune as a LoRA for the base model. |
Text encoder
qwen3vl_4b_bf16.safetensors from
Comfy-Org/Qwen3-VL. Load it
with Load CLIP, type krea2.
Using them
Text to image. Load Diffusion Model, Load CLIP (krea2), Load VAE (pixel_space), Empty
Latent Image at about 1 megapixel, and an empty negative prompt. res_multistep or euler,
simple, 50 steps, CFG 3.
Upscale x4. Base model plus the upscaler LoRA. Scale the image up 4x with bicubic, VAE Encode, multiply the latent by 2 (LatentMultiply), then one KSamplerAdvanced step: add_noise disable, steps 2, start at step 1, CFG 1, with ModelSamplingSD3 shift 1. Inputs of 256 to 384 px work best.
Depth. Scale the image to roughly 0.6 to 1 megapixel, VAE Encode, then one KSamplerAdvanced step with add_noise disabled and CFG 1. Small inputs (under about 600 px) come out blocky, so scale them up first. The output is relative depth: dark is near, bright is far.
How they were made
- bf16: a direct cast of the FP32 release.
- int8 convrot: the trunk MLP width (6826) is zero-padded to 6912 so every layer can be
quantized. The padding doesn't change the output. Quantized with comfy-model-tools'
quant_int8_convrot.py(mean weight error 0.85%). ComfyUI-Iris pads every checkpoint the same way when loading, so one LoRA fits all versions. - Depth: the released depth model takes an extra input channel that is always zero, plus a small output layer that turns three channels into one. Both are merged into the standard Iris layout without changing the result (maximum difference 0.00014 in FP32). The file is tagged so ComfyUI-Iris knows to output the depth map directly.
- Upscaler LoRA: the difference between the restoration fine-tune and the base model, compressed to rank 128 with SVD. Norms, biases and embeddings are stored exactly. Its output matches the full fine-tune at about 37 dB PSNR.
The conversion scripts are in the scripts/ folder of ComfyUI-Iris.
License
Apache-2.0, same as the original. The model and both fine-tunes are by Speridlabs (paper, code).
Model tree for nynxz/iris-3b
Base model
speridlabs/iris-3b