Attention kernel prebuilt wheels

Prebuilt attention-kernel wheels that Odyssey projects install through uv, built locally from pinned upstream source plus a small published patch. Hosted on the Hugging Face CDN so uv sync never has to compile CUDA.

Consuming

Reference a wheel by direct URL in [tool.uv.sources]:

[[tool.uv.sources.sageattention]]
url = "https://huggingface.co/amarkovich/attention-wheels/resolve/main/wheels/odyssey-acd75b5/sageattention-2.2.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl"
marker = "python_full_version == '3.10.*' and platform_machine == 'x86_64'"

The repo is public, so no HF_TOKEN is needed at install time.

Paths keep the build tag as a directory (wheels/<tag>/<filename>), so two builds that produce the same wheel filename never collapse into one. The local version segment (+cu130torch2.9cxx11abitrue) is set in each wheel's METADATA as well as its filename, so uv sees the installed version equal to the resolved one and does not reinstall on every run.

Contents

wheels/odyssey-acd75b5/ โ€” SageAttention for torch 2.9.1+cu130

Source: thu-ml/SageAttention @ d1a57a5 + sageattention-odyssey.patch (two commits: launch Sage2 kernels on torch's current CUDA stream, required for CUDA graphs; reproducible multi-arch packaging). Built with CUDA 13.0, torch 2.9.1+cu130 (C++11 ABI), CPython 3.10, x86_64.

file archs sha256 bytes
sageattention-2.2.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl sm80, sm89, sm90a, sm120a 8989b2bdd72b82d2ddb47aecedbad89c49a97ea9080ce3bc54e3bb6924922c8f 39,497,751
sageattn3-1.0.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl sm100a, sm120a f0fb1e7e22e6dc0232e42522d01260b5be5dd1d4592b28acb6935431b4e0cac7 2,764,183
sageattention-odyssey.patch source (git am onto d1a57a5, tree c1f23d3) dd86a4dfb3689f2cfe7bd4dd6ad4e78c38d1b1ec7f83ce66ed4ac235b0718156 21,585

Rebuilding

git clone https://github.com/thu-ml/SageAttention.git && cd SageAttention && git checkout d1a57a5
curl -sSfL https://huggingface.co/amarkovich/attention-wheels/resolve/main/wheels/odyssey-acd75b5/sageattention-odyssey.patch | git am
export CUDA_HOME=<cuda-13.0> PATH=<cuda-13.0>/bin:$PATH SAGEATTENTION_LOCAL_VERSION=cu130torch2.9cxx11abiTRUE
TORCH_CUDA_ARCH_LIST="8.0;8.9;9.0;12.0" EXT_PARALLEL=4 MAX_JOBS=8 python setup.py bdist_wheel
(cd sageattention3_blackwell && TORCH_CUDA_ARCH_LIST="10.0;12.0" MAX_JOBS=4 python setup.py bdist_wheel)

Run it with a Python that has torch==2.9.1+cu130, setuptools<75, wheel<0.44, packaging, and ninja installed.

Adding a wheel

Build from pinned source, upload under wheels/<tag>/ with the patch that produced it, verify the sha256 by downloading through the resolve/ URL (an HTTP 200 is not integrity evidence), and add a row above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support