InvokeAI SDXL-GGUF
Native SDXL GGUF-quantized UNet support for InvokeAI 6.x β self-describing GGUF files, a new model config + loader, and a CUDA-safe dequant dispatch. Install the model from a HuggingFace URL through the Model Manager, generate images with the UNet resident at ~3 GB instead of ~7 GB (measured end-to-end VRAM: 5.2 GB vs 15.4 GB).
GGUF file (Q8_0, 3.04 GB)
ββ quantized UNet weights (tensors, plain diffusers keys)
ββ self-describing header KVs
ββ sdxl.unet_config (UNet2DConditionModel architecture JSON)
ββ sdxl.source_checkpoint (path of origin checkpoint: TE + VAE)
ββ quantization.type (Q8_0)
β
βΌ
Main_GGUF_SDXL_Config ββclassifies by header only (no weight read)βββΊ
StableDiffusionXlGgufModel loader
ββ UNet β GGUF (quantized bytes stay resident, dequant-on-use)
ββ text encoders + VAE β any installed SDXL checkpoint (portable fallback)
Live demos
| ZeroGPU (Gradio) β UNet dequant-once, full fp16 speed on a shared RTX Pro 6000 | AEUPH/sdxl-gguf-q8 |
| InvokeAI UI (Docker) β full patched InvokeAI web UI booting the GGUF | AEUPH/invokeai-sdxl-gguf |
Benchmarks (RTX 3090 24 GB, 1024Γ1024, 30 steps, CFG 7, euler)
| model | format | wall | compute | VRAM after |
|---|---|---|---|---|
| SDXL ckpt (fp8 storage) | checkpoint | 11.18 s | 9.76 s | 15,412 MB |
| SDXL UNet Q8 GGUF | gguf_quantized | 17.52 s | 15.51 s | 5,172 MB |
VRAM drops 74 % (15.4 GB β 5.2 GB). Speed cost comes from the per-op dequant
fallback (every aten.linear dequantizes on the fly β no fused CUDA kernel);
see Roadmap.
Install (as a patch)
python apply_patches.py # auto-detects site-packages of this interpreter
# or
python apply_patches.py --site-packages /path/to/site-packages
The script is idempotent, backs originals up to patches/orig/, applies the
four diffs, py_compiles everything, and runs a registry smoke test
(ALL GREEN = loader registered). Then restart InvokeAI.
Patches:
| file | change |
|---|---|
configs/main.py |
Main_GGUF_SDXL_Config β classifies GGUFs by header KVs + add_embedding. tensor check |
configs/factory.py |
import + AnyModelConfig union entry |
load/model_loaders/stable_diffusion.py |
StableDiffusionXlGgufModel loader with portable source-checkpoint fallback |
quantization/gguf/ggml_tensor.py |
CUDA dequant fallback for aten.linear / conv2d / group_norm / layer_norm (upstream only mapped linear on MPS) |
Install the model (HuggingFace URL)
Once the patch is applied, Model Manager β Install from URL:
https://huggingface.co/AEUPH/sdxl-unet-q8-gguf/resolve/main/sdxl-unet-q8.gguf
The classifier reads only the GGUF header (no weight scan), matches the
sdxl.* KVs, and registers it as main.gguf_quantized.sdxl.normal.
Text encoders + VAE are resolved from the embedded sdxl.source_checkpoint
path, falling back to any installed SDXL single-file checkpoint when that
path doesn't exist (e.g. installs on other machines).
Convert your own checkpoint
python tools/sdxl_to_gguf.py --checkpoint /path/to/sdxl.safetensors --out sdxl-unet-q8.gguf
Quantizes every 2-D weight to Q8_0 (last-dim % 32 must hold; convs with
non-conforming shapes stay F16) and writes the sdxl.* self-description KVs.
Layout
patches/ 4 unified diffs + orig/ backups (generated against InvokeAI 6.14.1)
tools/ converter, patcher, bench harness, GGUF header dumper, remote-ops helpers
results/ bench JSONs + conversion/cluster logs
apply_patches.py idempotent patch installer + smoke test
Findings worth knowing
- Q8_0 write path:
gguf.quants.quantize()works on matrices, butGGUFWriter.add_tensor()rejects uint8 β passraw_dtype=qtso the writer accepts it and re-derives the dequantized shape. - Q4_K cannot be written by the installed gguf-py (
quantize_blocksraisesNotImplementedError); reading is fine. Use llama.cpp or upgrade gguf-py. - Weight-space similarity does not index tensors cross-family (cluster test
v1+v2: at/below control) β bundle/model management must key on name + role
metadata, which is exactly what the
sdxl.*KVs provide. - GGUF v1 files use u32 KV lengths (v2+ use u64) β header dumpers must branch.
Roadmap
- Row-block sidecar patching for LoRA on quantized modules
(
force_sidecar_patching=Trueβ direct in-place add would materialize a second fp16 copy). - Fused dequant+CUDA kernels (or cache the dequantized weights) to recover the ~1.6Γ speed gap.
- Deeper quant ladder: Q4_0 (write works today) β Q4_K via llama.cpp β AQLM/PVQ multi-vector β superposition/HDC bundling.
- Multi-component GGUF bundles (UNet + encoders + VAE + vision tower under
unet.*/text_encoder.*prefixes,bundle.componentsKV).
License
Apache-2.0 (same as InvokeAI). The model file on HuggingFace is released under the license of its origin SDXL checkpoint.
Model tree for AEUPH/invokeai-sdxl-gguf
Base model
stabilityai/stable-diffusion-xl-base-1.0