Buckets:
| <meta charset="utf-8" /><meta name="hf:doc:metadata" content="{"title":"Attention backends","local":"attention-backends","sections":[{"title":"Set a backend on the model","local":"set-a-backend-on-the-model","sections":[],"depth":2},{"title":"Try a backend temporarily","local":"try-a-backend-temporarily","sections":[],"depth":2},{"title":"Trusting remote kernels","local":"trusting-remote-kernels","sections":[],"depth":2},{"title":"Checks","local":"checks","sections":[],"depth":2},{"title":"Available backends","local":"available-backends","sections":[],"depth":2}],"depth":1}"/> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/entry/start.By0LmZnP.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/oRmjkTll.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/DK803DsY.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/entry/app.1jnSU0K4.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/DTwaC60R.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/BTASUwav.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/WKU8S240.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/DsnmJJEf.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/nodes/0.D3deY4PO.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/BzOvRKAw.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/nodes/281.CC4F4x1D.js" rel="modulepreload"> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/chunks/Bqo1LwON.js" rel="modulepreload"> | |
| <!--gxcprv--><meta name="hf:doc:metadata" content="{"title":"Attention backends","local":"attention-backends","sections":[{"title":"Set a backend on the model","local":"set-a-backend-on-the-model","sections":[],"depth":2},{"title":"Try a backend temporarily","local":"try-a-backend-temporarily","sections":[],"depth":2},{"title":"Trusting remote kernels","local":"trusting-remote-kernels","sections":[],"depth":2},{"title":"Checks","local":"checks","sections":[],"depth":2},{"title":"Available backends","local":"available-backends","sections":[],"depth":2}],"depth":1}"/><!----> | |
| <link href="/docs/diffusers/pr_14867/en/_app/immutable/assets/0.tn0RQdqM.css" rel="modulepreload"> <!--[--><!--[0--><!--[--><!--[0--><!--[--><p></p> <div class="items-center shrink-0 min-w-[100px] max-sm:min-w-[50px] justify-end ml-auto flex" style="float: right; margin-left: 10px; display: inline-flex; position: relative; z-index: 10;"><div class="inline-flex rounded-md max-sm:rounded-sm"><button class="inline-flex items-center gap-1 h-7 max-sm:h-7 px-2 max-sm:px-1.5 text-sm font-medium text-gray-800 border border-r-0 rounded-l-md max-sm:rounded-l-sm border-gray-200 bg-white hover:shadow-inner dark:border-gray-850 dark:bg-gray-950 dark:text-gray-200 dark:hover:bg-gray-800" aria-live="polite"><span class="inline-flex items-center justify-center rounded-md p-0.5 max-sm:p-0 hover:text-gray-800 dark:hover:text-gray-200"><svg class="sm:size-3.5 size-3" xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----></span> <span>Copy page</span></button> <button class="inline-flex items-center justify-center w-6 max-sm:w-5 h-7 max-sm:h-7 disabled:pointer-events-none text-sm text-gray-500 hover:text-gray-700 dark:hover:text-white rounded-r-md max-sm:rounded-r-sm border border-l transition border-gray-200 bg-white hover:shadow-inner dark:border-gray-850 dark:bg-gray-950 dark:text-gray-200 dark:hover:bg-gray-800" aria-haspopup="menu" aria-expanded="false" aria-label="Open copy menu"><svg class="transition-transform text-gray-400 overflow-visible sm:size-3.5 size-3 rotate-0" width="1em" height="1em" viewBox="0 0 12 7" fill="none" xmlns="http://www.w3.org/2000/svg"><path d="M1 1L6 6L11 1" stroke="currentColor"></path></svg><!----></button></div> <!--[-1--><!--]--></div><!----> <!--[0--><h1 class="relative group"><a id="attention-backends" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#attention-backends"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Attention backends</span></h1><!--]--><!----> <blockquote class="note"><p>The attention dispatcher is an experimental feature. Please open an issue if you have any feedback or encounter any problems.</p></blockquote> <p>Most Diffusers transformer models route attention through an <em>attention dispatcher</em> so you can switch optimized backends behind one API. The dispatcher manages registered implementations and exposes a unified call path for them. Some models, such as many autoencoders, don’t use the dispatcher because of their internals, and switching the backend has no effect on their attention layers.</p> <p>Refer to the table below for an overview of the available attention families and to the <a href="#available-backends">Available backends</a> section for a more complete list. The fastest backend depends on the model, GPU, dtype, and input shape.</p> <table><thead><tr><th>attention family</th><th>main feature</th></tr></thead><tbody><tr><td>FlashAttention</td><td>minimizes memory reads/writes through tiling and recomputation</td></tr><tr><td>AI Tensor Engine for ROCm</td><td>FlashAttention implementation optimized for AMD ROCm accelerators</td></tr><tr><td>SageAttention</td><td>quantizes attention to int8</td></tr><tr><td>FlexAttention</td><td>PyTorch FlexAttention</td></tr><tr><td>PyTorch native</td><td>built-in PyTorch implementation using <a href="./fp16#scaled-dot-product-attention">scaled_dot_product_attention</a></td></tr><tr><td>xFormers</td><td>memory-efficient attention with support for various attention kernels</td></tr></tbody></table> <p>Hub backends (the <code>*_hub</code> names) only need the <a href="https://github.com/huggingface/kernels" rel="nofollow">Kernels</a> library and download the kernel on first use. Other backends need their own package, such as <code>flash-attn</code> or <code>sageattention</code>. The <a href="#available-backends">Available backends</a> table lists the requirements Diffusers checks when you enable a backend.</p> <!--[1--><h2 class="relative group"><a id="set-a-backend-on-the-model" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#set-a-backend-on-the-model"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Set a backend on the model</span></h2><!--]--><!----> <p>The <a href="/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.set_attention_backend">set_attention_backend()</a> method walks the model’s attention layers and applies the chosen backend on each one. It also sets the dispatcher’s process-wide active backend to the same value.</p> <p><a href="/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.reset_attention_backend">reset_attention_backend()</a> clears the backend on attention layers only. It does not clear the process-wide active backend. For a temporary switch that restores the previous active backend on exit, use the <a href="#try-a-backend-temporarily">attention_backend</a> context manager.</p> <p>The example below enables <code>_flash_3_hub</code> (FlashAttention-3 from the Hub) with <code>device_map="cuda"</code>.</p> <div class="code-block relative "><div class="absolute top-2.5 right-4"><button class="inline-flex items-center relative text-sm focus:text-green-500 cursor-pointer focus:outline-none transition duration-200 ease-in-out opacity-0 mx-0.5 text-gray-600 " title="code excerpt" type="button"><svg xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----> <div class=" absolute pointer-events-none transition-opacity bg-black text-white py-1 px-2 leading-tight rounded font-normal shadow left-1/2 top-full transform -translate-x-1/2 translate-y-2 opacity-0 "><div class="absolute bottom-full left-1/2 transform -translate-x-1/2 w-0 h-0 border-black border-4 border-t-0" style="border-left-color: transparent; border-right-color: transparent;"></div> Copied</div><!----></button><!----></div> <pre class="language-py "><!----><span class="hljs-keyword">import</span> torch | |
| <span class="hljs-keyword">from</span> diffusers <span class="hljs-keyword">import</span> QwenImagePipeline | |
| pipeline = QwenImagePipeline.from_pretrained( | |
| <span class="hljs-string">"Qwen/Qwen-Image"</span>, dtype=torch.bfloat16, device_map=<span class="hljs-string">"cuda"</span> | |
| ) | |
| pipeline.transformer.set_attention_backend(<span class="hljs-string">"_flash_3_hub"</span>) | |
| prompt = <span class="hljs-string">""" | |
| cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California | |
| highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain | |
| """</span> | |
| pipeline(prompt).images[<span class="hljs-number">0</span>]<!----></pre></div><!----> <blockquote class="note"><p>The non-Hub FlashAttention-3 backends (<code>_flash_3</code>, <code>_flash_varlen_3</code>) require building FlashAttention-3 from source and will be deprecated soon. Use <code>_flash_3_hub</code> or <code>_flash_3_varlen_hub</code> instead.</p></blockquote> <!--[1--><h2 class="relative group"><a id="try-a-backend-temporarily" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#try-a-backend-temporarily"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Try a backend temporarily</span></h2><!--]--><!----> <p>The <code>attention_backend()</code> context manager sets the process-wide active backend for the duration of the block and restores the previous backend when the block exits. Use it to try a backend for one call without leaving a permanent backend applied from <a href="/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.set_attention_backend">set_attention_backend()</a>.</p> <div class="code-block relative "><div class="absolute top-2.5 right-4"><button class="inline-flex items-center relative text-sm focus:text-green-500 cursor-pointer focus:outline-none transition duration-200 ease-in-out opacity-0 mx-0.5 text-gray-600 " title="code excerpt" type="button"><svg xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----> <div class=" absolute pointer-events-none transition-opacity bg-black text-white py-1 px-2 leading-tight rounded font-normal shadow left-1/2 top-full transform -translate-x-1/2 translate-y-2 opacity-0 "><div class="absolute bottom-full left-1/2 transform -translate-x-1/2 w-0 h-0 border-black border-4 border-t-0" style="border-left-color: transparent; border-right-color: transparent;"></div> Copied</div><!----></button><!----></div> <pre class="language-py "><!----><span class="hljs-keyword">import</span> torch | |
| <span class="hljs-keyword">from</span> diffusers <span class="hljs-keyword">import</span> QwenImagePipeline, attention_backend | |
| pipeline = QwenImagePipeline.from_pretrained( | |
| <span class="hljs-string">"Qwen/Qwen-Image"</span>, dtype=torch.bfloat16, device_map=<span class="hljs-string">"cuda"</span> | |
| ) | |
| prompt = <span class="hljs-string">""" | |
| cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California | |
| highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain | |
| """</span> | |
| <span class="hljs-keyword">with</span> attention_backend(<span class="hljs-string">"_flash_3_hub"</span>): | |
| image = pipeline(prompt).images[<span class="hljs-number">0</span>]<!----></pre></div><!----> <blockquote class="tip"><p>Most attention backends work with <code>torch.compile</code>. Whether that speeds up your pipeline depends on the model and backend. See <a href="./fp16#torchcompile">Precision and compilation</a>.</p></blockquote> <!--[1--><h2 class="relative group"><a id="trusting-remote-kernels" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#trusting-remote-kernels"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Trusting remote kernels</span></h2><!--]--><!----> <p>Hub backends need the <a href="https://github.com/huggingface/kernels" rel="nofollow">Kernels</a> library first, and then Diffusers fetches the Hub kernel on first use.</p> <p>Hub attention backends download compute kernels with the <a href="https://github.com/huggingface/kernels" rel="nofollow">Kernels</a> library and run them locally. Most Hub attention names (<code>_flash_3_hub</code>, <code>flash_hub</code>, and the other <code>*_hub</code> backends) resolve to the <a href="https://huggingface.co/kernels-community" rel="nofollow">kernels-community</a> organization. That organization is a trusted publisher in Kernels, so Diffusers loads those attention kernels without setting <code>DIFFUSERS_TRUST_REMOTE_KERNELS</code>.</p> <blockquote class="note"><p>The SageAttention Hub backends (<code>sage_hub</code> and <code>sage_blackwell_hub</code>) load from the <a href="https://huggingface.co/SageAttention" rel="nofollow">SageAttention</a> organization instead. Set <code>DIFFUSERS_TRUST_REMOTE_KERNELS=true</code> to use them.</p></blockquote> <p>Other kernel-backed features such as <a href="../quantization/gguf">GGUF</a> and <a href="../quantization/nunchaku">Nunchaku Lite</a> can pull kernels from publishers outside kernels-community. Those paths stay blocked unless you opt in with <code>DIFFUSERS_TRUST_REMOTE_KERNELS</code>. When set, Diffusers forwards <code>trust_remote_code=True</code> to Kernels so untrusted publishers can load too.</p> <div class="code-block relative "><div class="absolute top-2.5 right-4"><button class="inline-flex items-center relative text-sm focus:text-green-500 cursor-pointer focus:outline-none transition duration-200 ease-in-out opacity-0 mx-0.5 text-gray-600 " title="code excerpt" type="button"><svg xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----> <div class=" absolute pointer-events-none transition-opacity bg-black text-white py-1 px-2 leading-tight rounded font-normal shadow left-1/2 top-full transform -translate-x-1/2 translate-y-2 opacity-0 "><div class="absolute bottom-full left-1/2 transform -translate-x-1/2 w-0 h-0 border-black border-4 border-t-0" style="border-left-color: transparent; border-right-color: transparent;"></div> Copied</div><!----></button><!----></div> <pre class="language-bash "><!----><span class="hljs-built_in">export</span> DIFFUSERS_TRUST_REMOTE_KERNELS=<span class="hljs-literal">true</span><!----></pre></div><!----> <p>Only enable this after inspecting the kernel repository. Without it, loading a kernel from an untrusted publisher raises an error. Diffusers performs this check itself, so it also applies to <code>kernels<0.14.0</code>, which predates the <code>trust_remote_code</code> argument. Setting <code>DIFFUSERS_DISABLE_REMOTE_CODE=true</code> disables remote code globally and takes precedence over <code>DIFFUSERS_TRUST_REMOTE_KERNELS</code>.</p> <!--[1--><h2 class="relative group"><a id="checks" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#checks"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Checks</span></h2><!--]--><!----> <p>The attention dispatcher can run debugging checks before each dispatched attention call. Which checks run depends on the constraints registered for the active backend.</p> <ol><li>Device checks verify that query, key, and value tensors live on the same device.</li> <li>Data type checks, where registered, confirm matching dtypes and often require <code>bfloat16</code> or <code>float16</code>.</li> <li>Shape checks validate tensor dimensions and prevent mixing attention masks with causal flags.</li></ol> <p>Enable checks with the <code>DIFFUSERS_ATTN_CHECKS</code> environment variable. Checks add overhead, so they are disabled by default.</p> <div class="code-block relative "><div class="absolute top-2.5 right-4"><button class="inline-flex items-center relative text-sm focus:text-green-500 cursor-pointer focus:outline-none transition duration-200 ease-in-out opacity-0 mx-0.5 text-gray-600 " title="code excerpt" type="button"><svg xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----> <div class=" absolute pointer-events-none transition-opacity bg-black text-white py-1 px-2 leading-tight rounded font-normal shadow left-1/2 top-full transform -translate-x-1/2 translate-y-2 opacity-0 "><div class="absolute bottom-full left-1/2 transform -translate-x-1/2 w-0 h-0 border-black border-4 border-t-0" style="border-left-color: transparent; border-right-color: transparent;"></div> Copied</div><!----></button><!----></div> <pre class="language-bash "><!----><span class="hljs-built_in">export</span> DIFFUSERS_ATTN_CHECKS=<span class="hljs-built_in">yes</span><!----></pre></div><!----> <p>With checks on, Diffusers runs those constraints before every dispatched attention call. The low-level example below calls <code>dispatch_attention_fn</code> directly. Pipeline inference does not need that import. It only needs the backend set via <a href="/docs/diffusers/pr_14867/en/api/models/overview#diffusers.ModelMixin.set_attention_backend">set_attention_backend()</a> or <code>attention_backend()</code>.</p> <div class="code-block relative "><div class="absolute top-2.5 right-4"><button class="inline-flex items-center relative text-sm focus:text-green-500 cursor-pointer focus:outline-none transition duration-200 ease-in-out opacity-0 mx-0.5 text-gray-600 " title="code excerpt" type="button"><svg xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M28,10V28H10V10H28m0-2H10a2,2,0,0,0-2,2V28a2,2,0,0,0,2,2H28a2,2,0,0,0,2-2V10a2,2,0,0,0-2-2Z" transform="translate(0)"></path><path d="M4,18H2V4A2,2,0,0,1,4,2H18V4H4Z" transform="translate(0)"></path><rect fill="none" width="32" height="32"></rect></svg><!----> <div class=" absolute pointer-events-none transition-opacity bg-black text-white py-1 px-2 leading-tight rounded font-normal shadow left-1/2 top-full transform -translate-x-1/2 translate-y-2 opacity-0 "><div class="absolute bottom-full left-1/2 transform -translate-x-1/2 w-0 h-0 border-black border-4 border-t-0" style="border-left-color: transparent; border-right-color: transparent;"></div> Copied</div><!----></button><!----></div> <pre class="language-py "><!----><span class="hljs-keyword">import</span> torch | |
| <span class="hljs-keyword">from</span> diffusers.models.attention_dispatch <span class="hljs-keyword">import</span> attention_backend, dispatch_attention_fn | |
| query = torch.randn(<span class="hljs-number">1</span>, <span class="hljs-number">10</span>, <span class="hljs-number">8</span>, <span class="hljs-number">64</span>, dtype=torch.bfloat16, device=<span class="hljs-string">"cuda"</span>) | |
| key = torch.randn(<span class="hljs-number">1</span>, <span class="hljs-number">10</span>, <span class="hljs-number">8</span>, <span class="hljs-number">64</span>, dtype=torch.bfloat16, device=<span class="hljs-string">"cuda"</span>) | |
| value = torch.randn(<span class="hljs-number">1</span>, <span class="hljs-number">10</span>, <span class="hljs-number">8</span>, <span class="hljs-number">64</span>, dtype=torch.bfloat16, device=<span class="hljs-string">"cuda"</span>) | |
| <span class="hljs-keyword">try</span>: | |
| <span class="hljs-keyword">with</span> attention_backend(<span class="hljs-string">"flash"</span>): | |
| output = dispatch_attention_fn(query, key, value) | |
| <span class="hljs-built_in">print</span>(<span class="hljs-string">"✓ Flash Attention works with checks enabled"</span>) | |
| <span class="hljs-keyword">except</span> Exception <span class="hljs-keyword">as</span> e: | |
| <span class="hljs-built_in">print</span>(<span class="hljs-string">f"✗ Flash Attention failed: <span class="hljs-subst">{e}</span>"</span>)<!----></pre></div><!----> <!--[1--><h2 class="relative group"><a id="available-backends" class="header-link block pr-1.5 text-lg no-hover:hidden with-hover:absolute with-hover:p-1.5 with-hover:opacity-0 with-hover:group-hover:opacity-100 with-hover:right-full" href="#available-backends"><span><svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 256 256"><path d="M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z" fill="currentColor"></path></svg><!----></span></a> <span>Available backends</span></h2><!--]--><!----> <p>Refer to the table below for a complete list of available attention backends and their variants. Diffusers checks package availability and version pins when you enable a backend.</p> <table><thead><tr><th>Backend Name</th><th>Family</th><th>Description</th><th>Prerequisite</th></tr></thead><tbody><tr><td><code>native</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>Default backend using PyTorch’s scaled_dot_product_attention</td><td>None</td></tr><tr><td><code>flex</code></td><td><a href="https://docs.pytorch.org/docs/stable/nn.attention.flex_attention.html#module-torch.nn.attention.flex_attention" rel="nofollow">FlexAttention</a></td><td>PyTorch FlexAttention</td><td><code>torch>=2.5.0</code></td></tr><tr><td><code>_native_cudnn</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>CuDNN-optimized attention</td><td>CUDA + CuDNN</td></tr><tr><td><code>_native_efficient</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>Memory-efficient attention</td><td>None beyond PyTorch</td></tr><tr><td><code>_native_flash</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>PyTorch’s FlashAttention</td><td>CUDA</td></tr><tr><td><code>_native_math</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>Math-based attention (fallback)</td><td>None</td></tr><tr><td><code>_native_npu</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>NPU-optimized attention</td><td><code>torch_npu</code></td></tr><tr><td><code>_native_xla</code></td><td><a href="https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend" rel="nofollow">PyTorch native</a></td><td>XLA-optimized attention</td><td><code>torch_xla>=2.2</code></td></tr><tr><td><code>flash</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>FlashAttention-2</td><td><code>flash-attn>=2.6.3</code></td></tr><tr><td><code>flash_hub</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>FlashAttention-2 from Hub kernels</td><td><code>kernels>=0.12</code></td></tr><tr><td><code>flash_varlen</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>Variable length FlashAttention</td><td><code>flash-attn>=2.6.3</code></td></tr><tr><td><code>flash_varlen_hub</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>Variable length FlashAttention from Hub kernels</td><td><code>kernels>=0.12</code></td></tr><tr><td><code>aiter_fa2_hub</code></td><td><a href="https://github.com/ROCm/aiter" rel="nofollow">AI Tensor Engine for ROCm</a></td><td>FlashAttention-2 for AMD ROCm from Hub kernels (<code>bfloat16</code>)</td><td><code>kernels>=0.12</code>, ROCm</td></tr><tr><td><code>flash_4_hub</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>FlashAttention-4 from Hub kernels</td><td><code>kernels>=0.12.3</code></td></tr><tr><td><code>_flash_3</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>FlashAttention-3 (local; deprecated soon)</td><td>Build FA3 from source</td></tr><tr><td><code>_flash_varlen_3</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>Variable length FlashAttention-3 (local; deprecated soon)</td><td>Build FA3 from source</td></tr><tr><td><code>_flash_3_hub</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>FlashAttention-3 from Hub kernels</td><td><code>kernels>=0.12</code></td></tr><tr><td><code>_flash_3_varlen_hub</code></td><td><a href="https://github.com/Dao-AILab/flash-attention" rel="nofollow">FlashAttention</a></td><td>Variable length FlashAttention-3 from Hub kernels</td><td><code>kernels>=0.12</code></td></tr><tr><td><code>sage</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>Quantized attention (INT8 QK)</td><td><code>sageattention>=2.1.1</code></td></tr><tr><td><code>sage_hub</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>Quantized attention (INT8 QK) from Hub kernels</td><td><code>kernels>=0.12</code>, <code>DIFFUSERS_TRUST_REMOTE_KERNELS=true</code></td></tr><tr><td><code>sage_blackwell_hub</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>SageAttention3 FP4 attention for SM120 Blackwell GPUs from Hub kernels</td><td><code>kernels>=0.12</code>, <code>DIFFUSERS_TRUST_REMOTE_KERNELS=true</code></td></tr><tr><td><code>sage_varlen</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>Variable length SageAttention</td><td><code>sageattention>=2.1.1</code></td></tr><tr><td><code>_sage_qk_int8_pv_fp8_cuda</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>INT8 QK + FP8 PV (CUDA)</td><td><code>sageattention>=2.1.1</code></td></tr><tr><td><code>_sage_qk_int8_pv_fp8_cuda_sm90</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>INT8 QK + FP8 PV (SM90)</td><td><code>sageattention>=2.1.1</code>; SM90</td></tr><tr><td><code>_sage_qk_int8_pv_fp16_cuda</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>INT8 QK + FP16 PV (CUDA)</td><td><code>sageattention>=2.1.1</code></td></tr><tr><td><code>_sage_qk_int8_pv_fp16_triton</code></td><td><a href="https://github.com/thu-ml/SageAttention" rel="nofollow">SageAttention</a></td><td>INT8 QK + FP16 PV (Triton)</td><td><code>sageattention>=2.1.1</code></td></tr><tr><td><code>xformers</code></td><td><a href="https://github.com/facebookresearch/xformers" rel="nofollow">xFormers</a></td><td>Memory-efficient attention</td><td><code>xformers>=0.0.29</code></td></tr></tbody></table> <a class="!text-gray-400 !no-underline text-sm flex items-center not-prose mt-4" href="https://github.com/huggingface/diffusers/blob/main/docs/source/en/optimization/attention_backends.md" target="_blank"><svg class="mr-1" xmlns="http://www.w3.org/2000/svg" aria-hidden="true" fill="currentColor" focusable="false" role="img" width="1em" height="1em" preserveAspectRatio="xMidYMid meet" viewBox="0 0 32 32"><path d="M31,16l-7,7l-1.41-1.41L28.17,16l-5.58-5.59L24,9l7,7z"></path><path d="M1,16l7-7l1.41,1.41L3.83,16l5.58,5.59L8,23l-7-7z"></path><path d="M12.419,25.484L17.639,6.552l1.932,0.518L14.351,26.002z"></path></svg><!----> <span><span class="underline">Update</span> on GitHub</span></a><!----> <p></p><!--]--><!----><!--]--><!--]--><!--]--> <!--[-1--><!--]--><!--]--> | |
| <script> | |
| { | |
| __sveltekit_asjqrd = { | |
| base: "/docs/diffusers/pr_14867/en", | |
| assets: "/docs/diffusers/pr_14867/en" | |
| }; | |
| const element = document.currentScript.parentElement; | |
| Promise.all([ | |
| import("/docs/diffusers/pr_14867/en/_app/immutable/entry/start.By0LmZnP.js"), | |
| import("/docs/diffusers/pr_14867/en/_app/immutable/entry/app.1jnSU0K4.js") | |
| ]).then(([kit, app]) => { | |
| kit.start(app, element, { | |
| node_ids: [0, 281], | |
| data: [null,null], | |
| form: null, | |
| error: null | |
| }); | |
| }); | |
| } | |
| </script> | |
Xet Storage Details
- Size:
- 36.1 kB
- Xet hash:
- 028e884dcc6aa8e7276c6325a47664df4a457227fec3e84c3aa031f5db2e82b3
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.