Nucleus-Image INT8 Convrot Experts + FP8 Text Encoder

Model overview

A quantized release of Nucleus-Image, a text-to-image diffusion model with a 17B-parameter mixture-of-experts transformer and approximately 2B active parameters per token.

This package combines INT8 ConvRot routed experts with the official Qwen3-VL-8B-Instruct-FP8. The DiT's other layers remain BF16. The original BF16 VAE, processor, tokenizer, and scheduler are included.

Total weight storage falls from 51.63 GB to 29.40 GB (43.1% smaller). Peak allocated VRAM during image generation falls from 25.80 GiB to 16.68 GiB (35.4% lower) in the comparison workflow described below.

Quantization

Component Format and scope Method / library
DiT routed experts INT8 weights, FP32 per-output-channel scales ConvRot through Comfy Kitchen 0.2.36
Remaining DiT layers Original BF16 weights, including attention, shared experts, routers, and the first three dense blocks Preserved from original
Text encoder Official FP8 E4M3 checkpoint Quantized and released by Qwen
VAE Original BF16 Preserved from original
Processor, tokenizer, scheduler Original configuration and vocabulary files Preserved from original

DiT conversion applies ConvRot to each expert matrix, then uses symmetric INT8 quantization with nearest rounding and a separate FP32 scale for every output row. Rotation groups are 256 for gate/up projections and 64 for down projections. The resulting 58 packed expert tensors cover 64 routed experts per MoE block; 552 other source tensors retain their original precision. Inference uses Comfy Kitchen's W8A8 ConvRot linear operations with BF16 outputs.

The text encoder is redistributed unchanged from Qwen's official FP8 release. Its embedding, normalization, vision, and output-head tensors retain their upstream precision.

Model size and VRAM

Weight-file sizes, including quantization scales and safetensors headers:

Component BF16 This release Reduction
DiT 33.84 GB 18.55 GB 45.2%
Text encoder 17.53 GB 10.59 GB 39.6%
VAE 0.25 GB 0.25 GB 0.0%
Total 51.63 GB 29.39 GB 43.1%

The storage table includes the complete text-encoder checkpoint. Configuration files and comparison images are additional.

Peak allocated VRAM BF16 DiT + BF16 TE INT8 experts DiT + FP8 TE Reduction
Image generation (denoising + VAE decode) 25.80 GiB 16.68 GiB 35.4%
Text encoding 11.63 GiB 6.43 GiB 44.7%

VRAM measurements use an 1024 Γ— 1024, 50 steps, CFG 4.0, an empty negative prompt, and one output image at a time. Generation values are the maximum across 12 prompts Γ— 2 seeds, with batched CFG, precomputed conditioning, PyTorch attention, CPU offload with a 4 GiB workspace reserve, and tiled VAE decode (512-pixel tiles, 64-pixel overlap).

Inference nodes

Dedicated ComfyUI inference nodes: coming soon.

The DiT uses the nucleus-int8-convrot checkpoint format and requires the dedicated loader. The package keeps the DiT, text encoder, and VAE in separate folders:

Path Contents
transformer/ INT8 experts-only DiT and BF16 remaining layers in 4 safetensors shards (up to 5 GB each), config, and quantization manifest
text_encoder/ Complete official Qwen3-VL-8B-Instruct-FP8 checkpoint and its configuration
vae/ Original BF16 VAE weights and config
processor/ Original Nucleus-Image processor, tokenizer, vocabulary, and chat template
scheduler/ Original flow-matching scheduler configuration
comparisons/ Original 1024-pixel comparison images and prompts
SHA256SUMS File-integrity checksums

Image comparisons

Left: BF16 DiT + BF16 text encoder. Right: INT8 experts-only DiT + FP8 text encoder. Both sides use the same prompt, seed, initial noise, and unchanged BF16 VAE. All examples below use seed 42, 50 steps / CFG 4.0 / 1024 Γ— 1024, with an empty negative prompt. Open an image for the full-resolution original. Prompts and image hashes are included in comparisons/prompts.json.

Landscape β€” Alpine lake at sunrise

BF16 DiT + BF16 TE INT8 experts DiT + FP8 TE
Landscape β€” Alpine lake at sunrise, BF16 Landscape β€” Alpine lake at sunrise, INT8 and FP8

Advertising β€” Luxury perfume campaign

BF16 DiT + BF16 TE INT8 experts DiT + FP8 TE
Advertising β€” Luxury perfume campaign, BF16 Advertising β€” Luxury perfume campaign, INT8 and FP8

Typography β€” English event poster

BF16 DiT + BF16 TE INT8 experts DiT + FP8 TE
Typography β€” English event poster, BF16 Typography β€” English event poster, INT8 and FP8

Fantasy β€” Dragon above the citadel

BF16 DiT + BF16 TE INT8 experts DiT + FP8 TE
Fantasy β€” Dragon above the citadel, BF16 Fantasy β€” Dragon above the citadel, INT8 and FP8

Quality and text rendering

Quantization can reduce image quality, including fine details, geometry, and lettering. Text-heavy prompts can show malformed or missing characters, with the severity varying by seed.

To fix text degradation caused by the FP8 text encoder, replace only the text encoder with Qwen3-VL-8B-Instruct (BF16). Keep the INT8 experts-only DiT, original processor, and VAE. The BF16 encoder is the same checkpoint used by the original Nucleus-Image pipeline.

The poster below uses the same prompt and seed as the typography comparison above. Switching only the text encoder back to BF16 restores the requested lettering:

INT8 DiT + FP8 text encoder INT8 DiT + BF16 text encoder
Poster with FP8 text encoder Poster after restoring the BF16 text encoder

Credits and license

The model weights are distributed under Apache 2.0. See LICENSE and NOTICE.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for nazunaex/Nucleus-Image-INT8-Convrot

Quantized
(2)
this model