Iris-3B W4A8 for Aikimi Forge Neo
Packed 4bit-weight / dynamic 8bit-activation derivatives of Iris-3B for text-to-image, relative depth, and restoration / 4x enlargement. A packed Qwen3-VL language decoder is included for text conditioning. Use the W4A8 model choice in Aikimi Forge Neo, alongside its original and weight-only INT8 choices. INT8 remains the default.
This custom checkpoint format (iris-convrot-w4a8-v1) requires Aikimi Forge Neo's packed loader. It uses Comfy Kitchen 0.2.33, ConvRot groups of 256, quantization groups of 16, symmetric learned codebooks, FP8 relative scales and FP32 channel scales. CUDA executes chunked INT4 decoding and INT8 GEMM with dynamic activation quantization. Internal BF16 fallback is rejected. Shared modulation cores, small projections, embeddings and normalization keep their source precision; this is a mixed-precision model, not an all-parameter 4bit model. Non-aligned projections are zero-padded for the kernel and cropped back to their original output dimensions.
The Qwen package contains only the language decoder and tokenizer used by Iris. Unused vision weights and the language-model output head are omitted. Depth and restoration use their original empty-prompt embeddings and do not load a text encoder.
Files and use
| Directory | Task | W4A8 checkpoint bytes |
|---|---|---|
| Root | Text-to-image model | 1,966,228,528 |
depth/ |
Relative depth model | 1,966,606,592 |
upscaler/ |
Restoration / 4x model | 1,966,228,528 |
text-encoder/ |
Qwen language conditioning | 2,826,604,984 |
Text-to-image weights total approximately 4.79 GB, compared with 12.11 GB for the existing INT8 body plus original separately downloaded Qwen checkpoint. These are disk sizes; they exclude configuration, tokenizer assets and the worker environment. The encoder disk saving includes omitting components that Iris does not use.
Update Aikimi Forge Neo to a version with the W4A8 option. On Windows with an Ampere-or-newer NVIDIA GPU and the CUDA 13 worker:
.\aikimi-iris-setup.bat --precision w4a8 --task generate
# All three task models and the packed text encoder:
.\aikimi-iris-setup.bat --precision w4a8 --task all
Or open Iris, select a task and W4A8, then choose モデルを準備. Aikimi Forge Neo retrieves a fixed validated Hub commit and checks file sizes and SHA-256 against Hub metadata and manifests. Users do not quantize locally. W4A8 text-to-image downloads its packed encoder from this repository and does not fetch the original BF16 Qwen files. The source inference code is pinned and unchanged; only upstream packaging makes training dependencies optional.
These files cannot be loaded directly by the original Iris loader or a Diffusers pipeline. The dedicated worker checks CUDA backend availability and rejects unsupported kernels with an operational error. Only RTX 3090 was physically tested.
Measured comparison
Tested on 2026-10-09, Windows, RTX 3090 24 GB, PyTorch 2.13.0+cu130. W4A8 ran with a 12 GiB PyTorch allocator limit, with internal BF16 fallback made fatal. This is not a physical 12 GB or 16 GB GPU acceptance test.
The same two prompts, seeds, inputs and task settings were used for normal, INT8 and W4A8. Text-to-image uses 1024x1024, 100 steps, CFG 3 and an empty negative prompt, after a 2-step loading warmup. Each condition was run once. Times include final GPU offload but exclude initial CPU model construction and downloads. Peaks below are PyTorch CUDA allocated memory, not total board usage.
| Image | Normal time / peak | INT8 time / peak | W4A8 time / peak |
|---|---|---|---|
| photo | 244.0s / 18.89 GiB | 244.1s / 10.68 GiB | 213.8s / 9.50 GiB |
| watercolor | 252.1s / 18.89 GiB | 254.5s / 10.68 GiB | 225.2s / 9.50 GiB |
The two generated scenes remain recognizable, while W4A8 changes expression, framing, lamp placement and fine detail more than INT8. These examples do not establish lossless quantization or a general quality guarantee.
All six full comparison cases completed with finite floating-point outputs. W4A8 relative depth agrees with normal at Pearson correlations 0.998806, 0.998867; this measures variant agreement, not ground-truth depth accuracy. Depth is relative, not metric distance. PNG visualizations and NPY arrays are exported by Aikimi Forge Neo.
Restoration used generated reference images downsampled and JPEG-compressed to 256x256 (quality 35) and 384x256 (quality 45), yielding 1024x1024 and 1536x1024 outputs. The latter exercises overlapping tiles. Floating-point finiteness was checked before integer image conversion. Visual inspection found similar structure with small texture, color and edge changes. Reconstructed details can differ from the source; no lossless restoration claim is made.
W4A8 depth/restoration allocated peaks were 5.61–5.70 GiB. Downstream had no warmup; first-case cold transfer and host/disk effects make its times unsuitable for a steady-state speed ranking. Maximum sampled whole-board usage across this run was 13.08 GiB at 250 ms intervals, including other processes; sampling can miss short spikes. No universal speed or lossless-quality claim is made. Prompts, conditions, exact measurements and comparisons are in benchmark.json.
Generation sizes follow the official trained buckets: 1024x1024, 1344x768, 1280x832, 1152x896, 896x1152, 832x1280, 768x1344. Restoration follows the official input budget (short edge up to 512, long edge up to 1024) and enlarges that processed input 4x.
Sources and license
- Iris weights: speridlabs/iris-3b.
- Iris inference: official source.
- Qwen language decoder / tokenizer: Qwen/Qwen3-VL-4B-Instruct.
- Kernel implementation: Comfy Kitchen, fixed runtime version 0.2.33.
Iris and Qwen model assets are Apache-2.0. LICENSE, original Iris NOTICE and derivative attribution are retained. This is Aikimi's independently converted derivative, not an official Speridlabs or Qwen W4A8 release.
Additional portrait anime tests
Two additional user-requested 832x1280 portrait cases use identical prompts and seeds across normal, INT8 and W4A8: an anime girl illustration, and an anime girl with a live-action-style photographic background. Each uses 100 steps, CFG 3 and an empty negative prompt, after a same-size 2-step warmup for each precision. Normal has the full allocator of the physical RTX3090 24GB; INT8 and W4A8 are limited to 12GiB. The allocation limit is not a physical 16GB acceptance test. Times include GPU offload and exclude initial CPU model construction. Each condition was measured once.
| Case | Normal time / allocated peak | INT8 time / allocated peak | W4A8 time / allocated peak |
|---|---|---|---|
| Anime illustration | 262.3s / 19.10 GiB | 281.2s / 10.89 GiB | 245.6s / 9.71 GiB |
| Anime character / live-action-style background | 267.6s / 19.10 GiB | 306.5s / 10.89 GiB | 244.1s / 9.71 GiB |
All six full-size outputs were visually reviewed. The anime illustration retains a blue-haired anime character and a painted cherry-blossom background across all three precisions. The mixed-media case retains a clearly drawn anime character against a photographic-looking rainy street across all three precisions. Normal and INT8 are close in composition; W4A8 changes facial expression, blouse details and umbrella/hand arrangement and has more visible grain in these outputs. No regular grid collapse was observed. The illustration hides the hands and crops the shoes at the lower edge, so it does not establish hand quality or complete full-body framing. The mixed-media images have imperfect umbrella handles, grip/hand details and compositing, and are not an anatomy or compositing quality guarantee.
Full-size model outputs:
| Case | Normal PNG | INT8 PNG | W4A8 PNG |
|---|---|---|---|
| Anime illustration | 832x1280 | 832x1280 | 832x1280 |
| Anime character / photographic background | 832x1280 | 832x1280 | 832x1280 |
Prompts, seeds, output hashes, measured allocator peaks and per-case whole-board memory samples are in portrait-benchmark.json. Visual observations are in portrait-visual-validation.json. These are generated images, not factual photographs. Style instructions and quality can vary by prompt and seed.
- Downloads last month
- 24
Model tree for Aikimi/iris-3b-w4a8
Base model
speridlabs/iris-3b




