Buckets:

|
download
raw
13.5 kB
# Precision and compilation
Lower precision and compilation are two ways to speed up Diffusers inference. Load weights in `bfloat16` or `float16`, then compile the denoiser (the UNet or the transformer) with `torch.compile` or regional compilation.
## Model data type
The precision and data type of the model weights affect inference speed because a higher precision requires more memory to load and more time to perform the computations. Diffusers loads model weights in `float32` when you omit `dtype`, so changing the data type is a simple way to quickly get faster inference.
`bfloat16` is similar to `float16` but it is more robust to numerical errors. Hardware support for `bfloat16` varies, but most modern GPUs are capable of supporting `bfloat16`.
```py
import torch
from diffusers import StableDiffusionXLPipeline
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.bfloat16
).to("cuda") # or "mps", "xpu", "cpu"
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
pipeline(prompt, num_inference_steps=30).images[0]
```
`float16` is similar to `bfloat16` but may be more prone to numerical errors.
```py
import torch
from diffusers import StableDiffusionXLPipeline
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16
).to("cuda") # or "mps", "xpu", "cpu"
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
pipeline(prompt, num_inference_steps=30).images[0]
```
[TensorFloat-32](https://blogs.nvidia.com/blog/2020/05/14/tensorfloat-32-precision-format/) (`tf32`) mode is supported on NVIDIA Ampere and newer GPUs and it computes the convolution and matrix multiplication operations in `tf32`. Storage and other operations are kept in `float32`. It speeds up operations that run in `float32`, so it helps most when the pipeline or parts of it stay in `float32`.
PyTorch only enables `tf32` mode for convolutions by default and you'll need to explicitly enable it for matrix multiplications.
```py
import torch
from diffusers import StableDiffusionXLPipeline
torch.backends.cuda.matmul.allow_tf32 = True
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float32
).to("cuda") # or "mps", "xpu", "cpu"
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
pipeline(prompt, num_inference_steps=30).images[0]
```
Refer to the [mixed precision training](https://huggingface.co/docs/transformers/en/perf_train_gpu_one#mixed-precision) docs for more details.
## Scaled dot product attention
[Scaled dot product attention (SDPA)](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) is the default attention on PyTorch 2.0 and later, through `AttnProcessor2_0` or the attention dispatcher's `native` backend. PyTorch picks an SDPA kernel for your hardware. For FlashAttention, SageAttention, xFormers, Hub kernels, and other backends, see [Attention backends](./attention_backends).
## torch.compile
[torch.compile](https://pytorch.org/tutorials/intermediate/torch_compile_tutorial.html) accelerates inference by compiling PyTorch code and operations into optimized kernels. You typically compile the denoiser that dominates runtime, either `pipeline.unet` or `pipeline.transformer`, and sometimes the VAE as well.
Enable the following compiler settings for maximum speed (refer to the [full list](https://github.com/pytorch/pytorch/blob/main/torch/_inductor/config.py) for more options).
```py
import torch
from diffusers import StableDiffusionXLPipeline
torch._inductor.config.conv_1x1_as_mm = True
torch._inductor.config.coordinate_descent_tuning = True
torch._inductor.config.epilogue_fusion = False
torch._inductor.config.coordinate_descent_check_all_directions = True
```
Load and compile the denoiser and VAE. Use `pipeline.unet` on UNet pipelines such as Stable Diffusion XL, or `pipeline.transformer` on Flux and DiT-style pipelines. There are several different modes you can choose from, but `"max-autotune"` searches for the fastest kernels and uses CUDA graphs. CUDA graphs reduce the overhead by launching multiple GPU operations through a single CPU operation.
Changing the memory layout to [channels_last](./memory#torchchannelslast) can also speed up convolution-heavy models like UNets and VAEs. Benchmark it first, because some models run slower.
```py
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16
).to("cuda") # or "mps", "xpu", "cpu"
pipeline.unet.to(memory_format=torch.channels_last)
pipeline.unet = torch.compile(
pipeline.unet, mode="max-autotune", fullgraph=True
)
pipeline.vae.to(memory_format=torch.channels_last)
pipeline.vae.decode = torch.compile(
pipeline.vae.decode,
mode="max-autotune",
fullgraph=True
)
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
pipeline(prompt, num_inference_steps=30).images[0]
```
Compilation is slow the first time, especially with `"max-autotune"`, because the compiler searches for kernels and keeps using the same compiled pipeline object afterward (see [Compile Time Caching](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_tutorial.html) and [Caching Configuration](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_configuration_tutorial.html) for how to configure the caching behavior). Calling the compiled pipeline on a different image size triggers recompilation.
### Dynamic shape compilation
`torch.compile` keeps track of input shapes and conditions, and if these are different, it recompiles the model. For example, if a model is compiled on a 1024x1024 resolution image and used on an image with a different resolution, it triggers recompilation.
To avoid recompilation, add `dynamic=True` to try and generate a more dynamic kernel to avoid recompilation when conditions change.
```diff
+ torch.fx.experimental._config.use_duck_shape = False
+ pipeline.unet = torch.compile(
pipeline.unet, fullgraph=True, dynamic=True
)
```
Specifying `use_duck_shape=False` instructs the compiler if it should use the same symbolic variable to represent input sizes that are the same. For more details, check out this [comment](https://github.com/huggingface/diffusers/pull/11327#discussion_r2047659790).
Not all models may benefit from dynamic compilation out of the box and may require changes. Refer to this [PR](https://github.com/huggingface/diffusers/pull/11297/) that improved the [AuraFlowPipeline](/docs/diffusers/pr_14867/en/api/pipelines/aura_flow#diffusers.AuraFlowPipeline) implementation to benefit from dynamic compilation.
Feel free to open an issue if dynamic compilation doesn't work as expected for a Diffusers model.
### Regional compilation
[Regional compilation](https://docs.pytorch.org/tutorials/recipes/regional_compilation.html) trims cold-start latency by only compiling the *small and frequently-repeated block(s)* of a model - typically a transformer layer - and enables reusing compiled artifacts for every subsequent occurrence.
For many diffusion architectures, this delivers comparable runtime speedups to full-graph compilation and can cut compile time by up to 8–10x.
Use the [compile_repeated_blocks()](/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.compile_repeated_blocks) method, a helper that wraps `torch.compile`, on the denoiser. Call it on `pipeline.unet` or `pipeline.transformer`, depending on the pipeline.
```py
import torch
from diffusers import StableDiffusionXLPipeline
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.float16,
).to("cuda") # or "mps", "xpu", "cpu"
pipeline.unet.compile_repeated_blocks(fullgraph=True)
```
To enable regional compilation for a new model, set `_repeated_blocks` to the class names of the repeated blocks (strings). For SDXL's UNet, use `BasicTransformerBlock`. Transformer denoisers list their own repeated block class names the same way.
```py
_repeated_blocks = ["BasicTransformerBlock"]
```
There is also a [compile_regions](https://github.com/huggingface/accelerate/blob/273799c85d849a1954a4f2e65767216eb37fa089/src/accelerate/utils/other.py#L78) method in [Accelerate](https://huggingface.co/docs/accelerate/index) that automatically selects candidate blocks in a model to compile. The remaining graph is compiled separately. This is useful for quick experiments because there aren't as many options for you to set which blocks to compile or adjust compilation flags.
```bash
pip install -U accelerate
```
```py
import torch
from diffusers import StableDiffusionXLPipeline
from accelerate.utils import compile_regions
pipeline = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16
).to("cuda") # or "mps", "xpu", "cpu"
pipeline.unet = compile_regions(pipeline.unet, mode="reduce-overhead", fullgraph=True)
```
[compile_repeated_blocks()](/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.compile_repeated_blocks) is intentionally explicit. List the blocks to repeat in `_repeated_blocks` and the helper only compiles those blocks. It offers predictable behavior and easy reasoning about cache reuse in one line of code.
### Graph breaks
Set `fullgraph=True` so torch.compile raises an error on a graph break instead of silently splitting the graph, which reduces the speedup.
### GPU sync
After each denoiser prediction, the pipeline calls the scheduler's `step` method, which indexes `sigmas`. Keeping `sigmas` on the GPU can force a CPU↔GPU sync on every step. That cost is easy to miss until the denoiser is compiled. Prefer leaving `sigmas` on the CPU (Diffusers schedulers such as Euler already do this).
### Benchmarks
Refer to the [diffusers/benchmarks](https://huggingface.co/datasets/diffusers/benchmarks) dataset to see inference latency and memory usage data for compiled pipelines.
The [diffusers-torchao](https://github.com/sayakpaul/diffusers-torchao#benchmarking-results) repository also contains benchmarking results for compiled versions of Flux and CogVideoX.
## Kernels
[Kernels](https://huggingface.co/docs/kernels/index) is a library for building, distributing, and loading optimized compute kernels on the [Hub](https://huggingface.co/kernels-community). It supports [attention](./attention_backends#set-a-backend-on-the-model) kernels and custom CUDA kernels for operations like RMSNorm, GEGLU, RoPE, and AdaLN.
The [Diffusers Pipeline Integration](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/references/diffusers-integration.md) guide shows how to integrate a kernel with the [add cuda-kernels](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/SKILL.md) skill. This skill enables an agent, like Claude or Codex, to write custom kernels targeted towards a specific model and your hardware. The [Custom kernels for all from Codex and Claude](https://huggingface.co/blog/custom-cuda-kernels-agent-skills) post has more detail.
For example, a custom RMSNorm kernel (generated by the `add cuda-kernels` skill) with [torch.compile](#torchcompile) speeds up LTX-Video generation 1.43x on an H100.
<iframe
src="https://huggingface.co/datasets/docs-benchmarks/kernel-ltx-video/embed/viewer/default/train"
frameborder="0"
width="100%"
height="560px"
>
## Dynamic quantization
Dynamic int8 activation quantization is available in [torchao](../quantization/torchao). Use [TorchAoConfig](/docs/diffusers/pr_14867/en/api/quantization#diffusers.TorchAoConfig) with a config such as [`Int8DynamicActivationInt8WeightConfig`](https://docs.pytorch.org/ao/stable/api_reference/generated/torchao.quantization.Int8DynamicActivationInt8WeightConfig.html), or torchao's `quantize_` API. Do not use the older `apply_dynamic_quant` helper.
## Fused projection matrices
> [!WARNING]
> `~DiffusionPipeline.fuse_qkv_projections` is experimental.
An input is projected into three subspaces, represented by the projection matrices Q, K, and V, in an attention block. These projections are typically calculated separately, but you can horizontally combine these into a single matrix and perform the projection in a single step. It increases the size of the matrix multiplications of the input projections and also improves the impact of quantization.
```py
pipeline.fuse_qkv_projections()
```
## Next steps
- Read the [Presenting Flux Fast: Making Flux go brrr on H100s](https://pytorch.org/blog/presenting-flux-fast-making-flux-go-brrr-on-h100s/) blog post to learn more about how you can combine all of these optimizations with [TorchInductor](https://docs.pytorch.org/docs/stable/torch.compiler.html) and [AOTInductor](https://docs.pytorch.org/docs/stable/torch.compiler_aot_inductor.html) for a ~2.5x speedup using recipes from [flux-fast](https://github.com/huggingface/flux-fast).
These recipes support AMD hardware and [Flux.1 Kontext Dev](https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev).
- Refer to the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post for maximizing performance with `torch.compile` for diffusion models.

Xet Storage Details

Size:
13.5 kB
·
Xet hash:
2d493d54c43423e246ff7455809a6d63ed25655d7339c8699fc3e16c0b7515ff

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.