Iris-3B W4A8 for Aikimi Forge Neo

Packed 4bit-weight / dynamic 8bit-activation derivatives of Iris-3B for text-to-image, relative depth, and restoration / 4x enlargement. A packed Qwen3-VL language decoder is included for text conditioning. Use the W4A8 model choice in Aikimi Forge Neo, alongside its original and weight-only INT8 choices. INT8 remains the default.

This custom checkpoint format (iris-convrot-w4a8-v1) requires Aikimi Forge Neo's packed loader. It uses Comfy Kitchen 0.2.33, ConvRot groups of 256, quantization groups of 16, symmetric learned codebooks, FP8 relative scales and FP32 channel scales. CUDA executes chunked INT4 decoding and INT8 GEMM with dynamic activation quantization. Internal BF16 fallback is rejected. Shared modulation cores, small projections, embeddings and normalization keep their source precision; this is a mixed-precision model, not an all-parameter 4bit model. Non-aligned projections are zero-padded for the kernel and cropped back to their original output dimensions.

The Qwen package contains only the language decoder and tokenizer used by Iris. Unused vision weights and the language-model output head are omitted. Depth and restoration use their original empty-prompt embeddings and do not load a text encoder.

Files and use

Directory Task W4A8 checkpoint bytes
Root Text-to-image model 1,966,228,528
depth/ Relative depth model 1,966,606,592
upscaler/ Restoration / 4x model 1,966,228,528
text-encoder/ Qwen language conditioning 2,826,604,984

Text-to-image weights total approximately 4.79 GB, compared with 12.11 GB for the existing INT8 body plus original separately downloaded Qwen checkpoint. These are disk sizes; they exclude configuration, tokenizer assets and the worker environment. The encoder disk saving includes omitting components that Iris does not use.

Update Aikimi Forge Neo to a version with the W4A8 option. On Windows with an Ampere-or-newer NVIDIA GPU and the CUDA 13 worker:

.\aikimi-iris-setup.bat --precision w4a8 --task generate
# All three task models and the packed text encoder:
.\aikimi-iris-setup.bat --precision w4a8 --task all

Or open Iris, select a task and W4A8, then choose モデルを準備. Aikimi Forge Neo retrieves a fixed validated Hub commit and checks file sizes and SHA-256 against Hub metadata and manifests. Users do not quantize locally. W4A8 text-to-image downloads its packed encoder from this repository and does not fetch the original BF16 Qwen files. The source inference code is pinned and unchanged; only upstream packaging makes training dependencies optional.

These files cannot be loaded directly by the original Iris loader or a Diffusers pipeline. The dedicated worker checks CUDA backend availability and rejects unsupported kernels with an operational error. Only RTX 3090 was physically tested.

Measured comparison

Tested on 2026-10-09, Windows, RTX 3090 24 GB, PyTorch 2.13.0+cu130. W4A8 ran with a 12 GiB PyTorch allocator limit, with internal BF16 fallback made fatal. This is not a physical 12 GB or 16 GB GPU acceptance test.

The same two prompts, seeds, inputs and task settings were used for normal, INT8 and W4A8. Text-to-image uses 1024x1024, 100 steps, CFG 3 and an empty negative prompt, after a 2-step loading warmup. Each condition was run once. Times include final GPU offload but exclude initial CPU model construction and downloads. Peaks below are PyTorch CUDA allocated memory, not total board usage.

Image Normal time / peak INT8 time / peak W4A8 time / peak
photo 244.0s / 18.89 GiB 244.1s / 10.68 GiB 213.8s / 9.50 GiB
watercolor 252.1s / 18.89 GiB 254.5s / 10.68 GiB 225.2s / 9.50 GiB

Photo: normal, INT8, W4A8

Watercolor: normal, INT8, W4A8

The two generated scenes remain recognizable, while W4A8 changes expression, framing, lamp placement and fine detail more than INT8. These examples do not establish lossless quantization or a general quality guarantee.

All six full comparison cases completed with finite floating-point outputs. W4A8 relative depth agrees with normal at Pearson correlations 0.998806, 0.998867; this measures variant agreement, not ground-truth depth accuracy. Depth is relative, not metric distance. PNG visualizations and NPY arrays are exported by Aikimi Forge Neo.

Relative depth: normal, INT8, W4A8

Restoration used generated reference images downsampled and JPEG-compressed to 256x256 (quality 35) and 384x256 (quality 45), yielding 1024x1024 and 1536x1024 outputs. The latter exercises overlapping tiles. Floating-point finiteness was checked before integer image conversion. Visual inspection found similar structure with small texture, color and edge changes. Reconstructed details can differ from the source; no lossless restoration claim is made.

Restoration: normal, INT8, W4A8

W4A8 depth/restoration allocated peaks were 5.61–5.70 GiB. Downstream had no warmup; first-case cold transfer and host/disk effects make its times unsuitable for a steady-state speed ranking. Maximum sampled whole-board usage across this run was 13.08 GiB at 250 ms intervals, including other processes; sampling can miss short spikes. No universal speed or lossless-quality claim is made. Prompts, conditions, exact measurements and comparisons are in benchmark.json.

Generation sizes follow the official trained buckets: 1024x1024, 1344x768, 1280x832, 1152x896, 896x1152, 832x1280, 768x1344. Restoration follows the official input budget (short edge up to 512, long edge up to 1024) and enlarges that processed input 4x.

Sources and license

Iris and Qwen model assets are Apache-2.0. LICENSE, original Iris NOTICE and derivative attribution are retained. This is Aikimi's independently converted derivative, not an official Speridlabs or Qwen W4A8 release.

Additional portrait anime tests

Two additional user-requested 832x1280 portrait cases use identical prompts and seeds across normal, INT8 and W4A8: an anime girl illustration, and an anime girl with a live-action-style photographic background. Each uses 100 steps, CFG 3 and an empty negative prompt, after a same-size 2-step warmup for each precision. Normal has the full allocator of the physical RTX3090 24GB; INT8 and W4A8 are limited to 12GiB. The allocation limit is not a physical 16GB acceptance test. Times include GPU offload and exclude initial CPU model construction. Each condition was measured once.

Case Normal time / allocated peak INT8 time / allocated peak W4A8 time / allocated peak
Anime illustration 262.3s / 19.10 GiB 281.2s / 10.89 GiB 245.6s / 9.71 GiB
Anime character / live-action-style background 267.6s / 19.10 GiB 306.5s / 10.89 GiB 244.1s / 9.71 GiB

Portrait anime illustration: normal, INT8, W4A8

Anime character and live-action-style background: normal, INT8, W4A8

All six full-size outputs were visually reviewed. The anime illustration retains a blue-haired anime character and a painted cherry-blossom background across all three precisions. The mixed-media case retains a clearly drawn anime character against a photographic-looking rainy street across all three precisions. Normal and INT8 are close in composition; W4A8 changes facial expression, blouse details and umbrella/hand arrangement and has more visible grain in these outputs. No regular grid collapse was observed. The illustration hides the hands and crops the shoes at the lower edge, so it does not establish hand quality or complete full-body framing. The mixed-media images have imperfect umbrella handles, grip/hand details and compositing, and are not an anatomy or compositing quality guarantee.

Full-size model outputs:

Case Normal PNG INT8 PNG W4A8 PNG
Anime illustration 832x1280 832x1280 832x1280
Anime character / photographic background 832x1280 832x1280 832x1280

Prompts, seeds, output hashes, measured allocator peaks and per-case whole-board memory samples are in portrait-benchmark.json. Visual observations are in portrait-visual-validation.json. These are generated images, not factual photographs. Style instructions and quality can vary by prompt and seed.

Downloads last month
24
Safetensors
Model size
2B params
Tensor type
F32
·
F8_E4M3
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aikimi/iris-3b-w4a8

Quantized
(3)
this model