Buckets:

hf-doc-build/doc-dev / diffusers /pr_14867 /en /stable_diffusion.md
HuggingFaceDocBuilder's picture
|
download
raw
4.39 kB
# Overview
Diffusion inference is computationally expensive, and you often run a [DiffusionPipeline](/docs/diffusers/pr_14867/en/api/pipelines/overview#diffusers.DiffusionPipeline) more than once before you like the result. This page provides an overview of the main Diffusers optimization techniques, what they do, and when to use them.
## Starter path
When the model fits on one GPU, start with this baseline load. Set `dtype` and place the pipeline on an accelerator. Reach for model CPU offload only when memory is tight. You could also speed up inference with fewer steps or a faster scheduler.
When you omit `dtype`, Diffusers loads components in `float32`. Pass `dtype=torch.bfloat16` (or `torch.float16` if bfloat16 is unsupported), then place the pipeline on an accelerator with `pipeline.to("cuda")`.
```py
import torch
from diffusers import DiffusionPipeline
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.to("cuda") # or "mps", "xpu"
prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
pipeline(prompt).images[0]
```
If the pipeline does not fit, or memory is tight, call [enable_model_cpu_offload()](/docs/diffusers/pr_14867/en/api/pipelines/overview#diffusers.DiffusionPipeline.enable_model_cpu_offload) instead of keeping everything on the GPU. It places the active model on the GPU and keeps the other components on the CPU.
Skip it when the model fits. Offloading is slower when you do not need it.
For more offloading options, see [Reduce memory usage](./optimization/memory#offloading).
```py
pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.enable_model_cpu_offload()
```
Lower latency with fewer `num_inference_steps` or a faster scheduler such as [DPMSolverMultistepScheduler](/docs/diffusers/pr_14867/en/api/schedulers/multistep_dpm_solver#diffusers.DPMSolverMultistepScheduler). That usually speeds up generation but can reduce image quality versus a slower, higher-quality scheduler. See [Optimization techniques](#optimization-techniques) below for more speed techniques.
```py
import time
from diffusers import DPMSolverMultistepScheduler
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config)
start_time = time.perf_counter()
image = pipeline(prompt, num_inference_steps=25).images[0]
end_time = time.perf_counter()
print(f"Image generation took {end_time - start_time:.3f} seconds")
```
## Optimization techniques
When the starter path is not enough, use these techniques. If inference is too slow, start with `torch.compile`, caching, or attention backends. If you are out of memory, start with offloading or quantization.
Faster inference:
- [torch.compile](./optimization/fp16#torchcompile) — Compile the UNet, transformer, or VAE into optimized kernels.
- [Regional compilation](./optimization/fp16#regional-compilation) — Compile repeated blocks to cut `torch.compile` cold-start latency and reuse compiled artifacts.
- [Caching](./optimization/cache) — Reuse intermediates across denoising steps when you want more speed and can spend memory.
- [Attention backends](./optimization/attention_backends) — Swap Diffusers attention implementations through a unified API when attention is the bottleneck.
- [Kernels](./optimization/fp16#kernels) — Load optimized Hub compute kernels (attention and custom CUDA ops such as RMSNorm or RoPE) when you need hardware-specific speedups beyond stock PyTorch.
Less memory:
- [Offloading](./optimization/memory#offloading) — Move inactive models or layers to the CPU with CPU, model, or group offloading.
- [Quantization](./quantization/overview) — Load smaller weights to cut memory (some backends also speed up inference). [GGUF](./quantization/gguf) is a common starting point.
- [VAE slicing](./optimization/memory#vae-slicing) and [VAE tiling](./optimization/memory#vae-tiling) — Decode large batches or high-resolution images in pieces to lower peak memory.
Both:
- [Quantize, compile, and offload](./optimization/speed-memory-optims) — Combine quantization, `torch.compile`, and offloading when one technique is not enough.

Xet Storage Details

Size:
4.39 kB
·
Xet hash:
84bc1cec4f7fad7f9f6102be870ff97d7f6776abc4bd47b4f2e88fe04f913a7e

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.