Iris-3B INT8 for Aikimi Forge Neo

Rowwise symmetric, weight-only INT8 derivatives of Speridlabs' Iris-3B image generation, relative depth, and restoration models. Use the Iris tab in Aikimi Forge Neo. The integration can also run the original, unquantized weights.

This is a custom checkpoint format (iris-rowwise-int8-v1) for Aikimi Forge Neo's Iris loader. It stores selected Linear weights as INT8 with per-output FP32 scales, then computes Linear operations in BF16 under CUDA autocast. It does not use an integer GEMM kernel. Shared modulation cores and other parameters retain their original format. The Qwen3-VL text encoder is unchanged and downloaded separately for image generation; depth and restoration do not load it.

Files and use

Directory Task INT8 model bytes Original model bytes
Root Text-to-image 3,233,164,048 11,950,112,680
depth/ Relative depth 3,233,825,264 11,952,738,184
upscaler/ Restoration and 4x enlargement 3,233,164,048 11,950,112,680

Each model includes its original configuration and a manifest.json recording the source revision, SHA-256, conversion format, quantized Linear count, and output file sizes and SHA-256. Downstream models also include their original empty-prompt embedding. Model weights alone are about 73% smaller. Text-to-image still needs the separate Qwen3-VL-4B-Instruct encoder (about 8.88 GB of weights).

On Windows with an NVIDIA GPU, update Aikimi Forge Neo and run:

.\aikimi-iris-setup.bat

Or select Iris, a task, and INT8 in the UI, then choose モデルを準備. Only the selected task's weights are downloaded. To prepare all three:

.\venv\Scripts\python.exe tools/setup_iris.py --precision int8 --task all

Aikimi Forge Neo downloads a fixed, validated Hub commit and checks every weight file against size and SHA-256. Users do not need to quantize locally. The dedicated worker environment is separate from the main UI environment. Source inference code is pinned; only its packaging definition moves training dependencies into an optional extra.

The format requires Aikimi Forge Neo's quantization loader. Loading these files directly with the original upstream loader or a Diffusers pipeline is not supported.

Text-to-image size choices follow the upstream demo's buckets with substantial training mass: 1024x1024 (default), 1344x768, 1280x832, 1152x896, 896x1152, 832x1280, and 768x1344. Additional INT8 UI tests at 256x256 produced collapsed grid-like output and at 512x512 visible grid artifacts, so these sizes are excluded from the integration. The paired comparison images below use 1024x1024.

Measured comparisons

Tested on 2026-10-09 with one NVIDIA RTX 3090 24 GB, Windows, Python 3.13.14, PyTorch 2.13.0+cu130. Each pair uses identical input and conditions. Text-to-image uses 1024x1024, 100 steps, CFG 3, an empty negative prompt, and fixed seeds. Cold-load smoke tests and initial CPU model construction are excluded from timing; inference wall time includes returning model weights to CPU. Peaks are PyTorch CUDA allocated memory, not total board usage. Each condition was measured once.

Text-to-image's measured peak decreased from approximately 18.9 GiB to 10.7 GiB. Generation times were similar, with INT8 slightly slower in these samples. Both normal and INT8 completed at 1024x1024 on this 24 GB card. No physical 12 GB or 16 GB card was tested.

Photo comparison

Watercolor comparison

Both generated pairs preserved the scene and subject in visual inspection, with small changes in fine detail. They are not pixel-identical, and these examples do not establish lossless quantization or a general quality guarantee. Prompts, seeds, raw measurements, and paired differences are in benchmark.json.

Depth is relative depth, not metric distance. The two 1024x1024 comparisons have finite depth arrays and normal/INT8 Pearson correlations of 0.999908 and 0.999945; this measures agreement between variants, not ground-truth depth accuracy. PNG visualization and NPY arrays are exported.

Relative depth comparison

Restoration inputs were deliberately downsampled and JPEG-compressed generated references: 256x256 quality 35 to 1024x1024 (one tile), and 384x256 quality 45 to 1536x1024 (two overlapping tiles). Both precisions completed both paths, and floating-point output was checked for finiteness before integer image conversion. Visual inspection found similar structure and detail in these pairs, with small changes. Restoration may reconstruct detail differently from the reference.

Restoration comparison

All twelve paired CUDA cases completed. Depth/restoration peak allocated memory was approximately 15.0–15.1 GiB normal and 6.8–6.9 GiB INT8. These tasks had no pre-run warmup; the first photo includes cold GPU transfer and host memory/disk effects, so their times must not be compared as a steady-state speed benchmark. See the methodology in benchmark.json.

Restoration follows the upstream input budget (short edge up to 512, long edge up to 1024) and enlarges that processed input 4x, using overlapping tiles where required.

Source and license

Apache-2.0. Original LICENSE and NOTICE are retained. Aikimi's modification is the weight-only conversion and Aikimi Forge Neo integration; this is an independently converted derivative, not an official Speridlabs INT8 release.

Additional portrait anime tests

Two additional user-requested 832x1280 portrait cases use identical prompts and seeds across normal, INT8 and W4A8: an anime girl illustration, and an anime girl with a live-action-style photographic background. Each uses 100 steps, CFG 3 and an empty negative prompt, after a same-size 2-step warmup for each precision. Normal has the full allocator of the physical RTX3090 24GB; INT8 and W4A8 are limited to 12GiB. The allocation limit is not a physical 16GB acceptance test. Times include GPU offload and exclude initial CPU model construction. Each condition was measured once.

Case Normal time / allocated peak INT8 time / allocated peak W4A8 time / allocated peak
Anime illustration 262.3s / 19.10 GiB 281.2s / 10.89 GiB 245.6s / 9.71 GiB
Anime character / live-action-style background 267.6s / 19.10 GiB 306.5s / 10.89 GiB 244.1s / 9.71 GiB

Portrait anime illustration: normal, INT8, W4A8

Anime character and live-action-style background: normal, INT8, W4A8

All six full-size outputs were visually reviewed. The anime illustration retains a blue-haired anime character and a painted cherry-blossom background across all three precisions. The mixed-media case retains a clearly drawn anime character against a photographic-looking rainy street across all three precisions. Normal and INT8 are close in composition; W4A8 changes facial expression, blouse details and umbrella/hand arrangement and has more visible grain in these outputs. No regular grid collapse was observed. The illustration hides the hands and crops the shoes at the lower edge, so it does not establish hand quality or complete full-body framing. The mixed-media images have imperfect umbrella handles, grip/hand details and compositing, and are not an anatomy or compositing quality guarantee.

Full-size model outputs:

Case Normal PNG INT8 PNG W4A8 PNG
Anime illustration 832x1280 832x1280 832x1280
Anime character / photographic background 832x1280 832x1280 832x1280

Prompts, seeds, output hashes, measured allocator peaks and per-case whole-board memory samples are in portrait-benchmark.json. Visual observations are in portrait-visual-validation.json. These are generated images, not factual photographs. Style instructions and quality can vary by prompt and seed.

Downloads last month
23
Safetensors
Model size
3B params
Tensor type
F32
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aikimi/iris-3b-int8

Quantized
(4)
this model