Instructions to use QuantFunc/Krea-2-QuantFunc-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use QuantFunc/Krea-2-QuantFunc-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("QuantFunc/Krea-2-QuantFunc-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("QuantFunc/Krea-2-QuantFunc-4bit", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]๐ Website | ๐ GitHub | ๐ค Hugging Face | ๐ค ModelScope | ๐ฎ Discord
Krea-2-QuantFunc-4bit
4x compression, quality held.
QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.
Showcase
All images below were generated by Krea-2-QuantFunc-4bit.
Fast denoising, fast end-to-end too
High-VRAM
On an RTX 4090, same workflow and generation settings:
| Stage | QuantFunc INT4 | FP8 | Speedup |
|---|---|---|---|
| Denoising | 1.6s | 5.0s | 3.13x |
| End-to-end | 2.5s | 6.5s | 2.6x |
End-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.
Low-VRAM
On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup โ up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.
Swap one loader, keep your workflow
- Install or update ComfyUI-QuantFunc.
- Download the r128 or r32 weights.
- Swap your model loader for the QuantFunc loader and pick the matching weight file.
Every other node, connection and generation parameter stays as-is.
Krea-2-Turbo is a distilled turbo model โ use fewer sampling steps and turn CFG off (guidance 1.0).
RTX 20-series through GB300, one build covers it all
Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.
Choose a model
| Variant | File | Size | Use case |
|---|---|---|---|
| r128 | krea2-turbo-quantfunc-int4-r128.safetensors |
~8.3 GB | Quality-first, recommended |
| r32 | krea2-turbo-quantfunc-int4-r32.safetensors |
~7.8 GB | Size/VRAM-first |
Loading
Weights use QuantFunc's own safetensors format โ load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ see HF's download-stats docs โ not because the files load via the diffusers package.)
Source & license
This repository distributes derived quantized weights produced from:
- Base model: krea/Krea-2-Turbo
- Pretrained base: krea/Krea-2-Raw
Weights follow the Krea 2 Community License โ please read the original model's license terms before use.
Community
- ๐ QuantFunc website
- ๐ ComfyUI-QuantFunc on GitHub
- ๐ฎ Discord
- Downloads last month
- 4,328



