How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("QuantFunc/Krea-2-QuantFunc-4bit", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

๐ŸŒ Website  |  ๐Ÿ™ GitHub  |  ๐Ÿค— Hugging Face  |  ๐Ÿค– ModelScope  |  ๐ŸŽฎ Discord

Krea-2-QuantFunc-4bit

4x compression, quality held.

QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.

Showcase

All images below were generated by Krea-2-QuantFunc-4bit.

Portrait photography Sci-fi scene
Portrait Sci-fi
Impasto art Watercolor illustration
Impasto Watercolor

Fast denoising, fast end-to-end too

High-VRAM

On an RTX 4090, same workflow and generation settings:

Stage QuantFunc INT4 FP8 Speedup
Denoising 1.6s 5.0s 3.13x
End-to-end 2.5s 6.5s 2.6x

End-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.

Low-VRAM

On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup โ€” up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.

Swap one loader, keep your workflow

  1. Install or update ComfyUI-QuantFunc.
  2. Download the r128 or r32 weights.
  3. Swap your model loader for the QuantFunc loader and pick the matching weight file.

Every other node, connection and generation parameter stays as-is.

Krea-2-Turbo is a distilled turbo model โ€” use fewer sampling steps and turn CFG off (guidance 1.0).

RTX 20-series through GB300, one build covers it all

Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

Choose a model

Variant File Size Use case
r128 krea2-turbo-quantfunc-int4-r128.safetensors ~8.3 GB Quality-first, recommended
r32 krea2-turbo-quantfunc-int4-r32.safetensors ~7.8 GB Size/VRAM-first

Loading

Weights use QuantFunc's own safetensors format โ€” load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ€” see HF's download-stats docs โ€” not because the files load via the diffusers package.)

Source & license

This repository distributes derived quantized weights produced from:

Weights follow the Krea 2 Community License โ€” please read the original model's license terms before use.

Community

Downloads last month
4,328
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for QuantFunc/Krea-2-QuantFunc-4bit

Base model

krea/Krea-2-Raw
Quantized
(45)
this model