Buckets:
| # Overview | |
| Diffusion inference is computationally expensive, and you often run a [DiffusionPipeline](/docs/diffusers/pr_14867/en/api/pipelines/overview#diffusers.DiffusionPipeline) more than once before you like the result. This page provides an overview of the main Diffusers optimization techniques, what they do, and when to use them. | |
| ## Starter path | |
| When the model fits on one GPU, start with this baseline load. Set `dtype` and place the pipeline on an accelerator. Reach for model CPU offload only when memory is tight. You could also speed up inference with fewer steps or a faster scheduler. | |
| When you omit `dtype`, Diffusers loads components in `float32`. Pass `dtype=torch.bfloat16` (or `torch.float16` if bfloat16 is unsupported), then place the pipeline on an accelerator with `pipeline.to("cuda")`. | |
| ```py | |
| import torch | |
| from diffusers import DiffusionPipeline | |
| pipeline = DiffusionPipeline.from_pretrained( | |
| "stabilityai/stable-diffusion-xl-base-1.0", | |
| dtype=torch.bfloat16, | |
| ) | |
| pipeline.to("cuda") # or "mps", "xpu" | |
| prompt = """ | |
| cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California | |
| highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain | |
| """ | |
| pipeline(prompt).images[0] | |
| ``` | |
| If the pipeline does not fit, or memory is tight, call [enable_model_cpu_offload()](/docs/diffusers/pr_14867/en/api/pipelines/overview#diffusers.DiffusionPipeline.enable_model_cpu_offload) instead of keeping everything on the GPU. It places the active model on the GPU and keeps the other components on the CPU. | |
| Skip it when the model fits. Offloading is slower when you do not need it. | |
| For more offloading options, see [Reduce memory usage](./optimization/memory#offloading). | |
| ```py | |
| pipeline = DiffusionPipeline.from_pretrained( | |
| "stabilityai/stable-diffusion-xl-base-1.0", | |
| dtype=torch.bfloat16, | |
| ) | |
| pipeline.enable_model_cpu_offload() | |
| ``` | |
| Lower latency with fewer `num_inference_steps` or a faster scheduler such as [DPMSolverMultistepScheduler](/docs/diffusers/pr_14867/en/api/schedulers/multistep_dpm_solver#diffusers.DPMSolverMultistepScheduler). That usually speeds up generation but can reduce image quality versus a slower, higher-quality scheduler. See [Optimization techniques](#optimization-techniques) below for more speed techniques. | |
| ```py | |
| import time | |
| from diffusers import DPMSolverMultistepScheduler | |
| pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config) | |
| start_time = time.perf_counter() | |
| image = pipeline(prompt, num_inference_steps=25).images[0] | |
| end_time = time.perf_counter() | |
| print(f"Image generation took {end_time - start_time:.3f} seconds") | |
| ``` | |
| ## Optimization techniques | |
| When the starter path is not enough, use these techniques. If inference is too slow, start with `torch.compile`, caching, or attention backends. If you are out of memory, start with offloading or quantization. | |
| Faster inference: | |
| - [torch.compile](./optimization/fp16#torchcompile) — Compile the UNet, transformer, or VAE into optimized kernels. | |
| - [Regional compilation](./optimization/fp16#regional-compilation) — Compile repeated blocks to cut `torch.compile` cold-start latency and reuse compiled artifacts. | |
| - [Caching](./optimization/cache) — Reuse intermediates across denoising steps when you want more speed and can spend memory. | |
| - [Attention backends](./optimization/attention_backends) — Swap Diffusers attention implementations through a unified API when attention is the bottleneck. | |
| - [Kernels](./optimization/fp16#kernels) — Load optimized Hub compute kernels (attention and custom CUDA ops such as RMSNorm or RoPE) when you need hardware-specific speedups beyond stock PyTorch. | |
| Less memory: | |
| - [Offloading](./optimization/memory#offloading) — Move inactive models or layers to the CPU with CPU, model, or group offloading. | |
| - [Quantization](./quantization/overview) — Load smaller weights to cut memory (some backends also speed up inference). [GGUF](./quantization/gguf) is a common starting point. | |
| - [VAE slicing](./optimization/memory#vae-slicing) and [VAE tiling](./optimization/memory#vae-tiling) — Decode large batches or high-resolution images in pieces to lower peak memory. | |
| Both: | |
| - [Quantize, compile, and offload](./optimization/speed-memory-optims) — Combine quantization, `torch.compile`, and offloading when one technique is not enough. | |
Xet Storage Details
- Size:
- 4.39 kB
- Xet hash:
- 84bc1cec4f7fad7f9f6102be870ff97d7f6776abc4bd47b4f2e88fe04f913a7e
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.