Buckets:
Overview
Diffusion inference is computationally expensive, and you often run a DiffusionPipeline more than once before you like the result. This page provides an overview of the main Diffusers optimization techniques, what they do, and when to use them.
Starter path
When the model fits on one GPU, start with this baseline load. Set dtype and place the pipeline on an accelerator. Reach for model CPU offload only when memory is tight. You could also speed up inference with fewer steps or a faster scheduler.
When you omit dtype, Diffusers loads components in float32. Pass dtype=torch.bfloat16 (or torch.float16 if bfloat16 is unsupported), then place the pipeline on an accelerator with pipeline.to("cuda").
import torch
from diffusers import DiffusionPipeline
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.to("cuda") # or "mps", "xpu"
prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
pipeline(prompt).images[0]
If the pipeline does not fit, or memory is tight, call enable_model_cpu_offload() instead of keeping everything on the GPU. It places the active model on the GPU and keeps the other components on the CPU.
Skip it when the model fits. Offloading is slower when you do not need it.
For more offloading options, see Memory and offloading.
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.enable_model_cpu_offload()
Lower latency with fewer num_inference_steps or a faster scheduler such as DPMSolverMultistepScheduler. That usually speeds up generation but can reduce image quality versus a slower, higher-quality scheduler. See Precision and compilation for more speed techniques.
import time
from diffusers import DPMSolverMultistepScheduler
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config)
start_time = time.perf_counter()
image = pipeline(prompt, num_inference_steps=25).images[0]
end_time = time.perf_counter()
print(f"Image generation took {end_time - start_time:.3f} seconds")
Optimization techniques
When the starter path is not enough, use these techniques. If you are out of memory, start with offloading or quantization. If inference is too slow, start with caching, attention backends, torch.compile, or regional compilation.
- Caching — Reuse intermediates across denoising steps when you want more speed and can spend memory.
- Attention backends — Swap Diffusers attention implementations through a unified API when attention is the bottleneck.
- Quantization — Load smaller weights to cut memory (and often speed up inference). GGUF is a common starting point.
- Regional compilation — Compile repeated blocks to cut
torch.compilecold-start latency and reuse compiled artifacts. - torch.compile — Compile the UNet, transformer, or VAE into optimized kernels.
- Kernels — Load optimized Hub compute kernels (attention and custom CUDA ops such as RMSNorm or RoPE) when you need hardware-specific speedups beyond stock PyTorch.
- Offloading — Move inactive models or layers to the CPU with CPU, model, or group offloading.
- Quantize, compile, and offload — Combine quantization,
torch.compile, and offloading when one technique is not enough.
Xet Storage Details
- Size:
- 4.17 kB
- Xet hash:
- 981f73cb66c07a373a0e4d8f34c166e99f470c5e678f21a1415b04dc89718e48
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.