Attention kernel prebuilt wheels
Prebuilt attention-kernel wheels that Odyssey projects install through uv, built locally from pinned upstream source plus a small published patch.
Hosted on the Hugging Face CDN so uv sync never has to compile CUDA.
Consuming
Reference a wheel by direct URL in [tool.uv.sources]:
[[tool.uv.sources.sageattention]]
url = "https://huggingface.co/amarkovich/attention-wheels/resolve/main/wheels/odyssey-acd75b5/sageattention-2.2.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl"
marker = "python_full_version == '3.10.*' and platform_machine == 'x86_64'"
The repo is public, so no HF_TOKEN is needed at install time.
Paths keep the build tag as a directory (wheels/<tag>/<filename>), so two builds that produce the same wheel filename never collapse into one.
The local version segment (+cu130torch2.9cxx11abitrue) is set in each wheel's METADATA as well as its filename, so uv sees the installed version equal to the resolved one and does not reinstall on every run.
Contents
wheels/odyssey-acd75b5/ โ SageAttention for torch 2.9.1+cu130
Source: thu-ml/SageAttention @ d1a57a5 + sageattention-odyssey.patch (two commits: launch Sage2 kernels on torch's current CUDA stream, required for CUDA graphs; reproducible multi-arch packaging).
Built with CUDA 13.0, torch 2.9.1+cu130 (C++11 ABI), CPython 3.10, x86_64.
| file | archs | sha256 | bytes |
|---|---|---|---|
| sageattention-2.2.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl | sm80, sm89, sm90a, sm120a | 8989b2bdd72b82d2ddb47aecedbad89c49a97ea9080ce3bc54e3bb6924922c8f | 39,497,751 |
| sageattn3-1.0.0+cu130torch2.9cxx11abitrue-cp310-cp310-linux_x86_64.whl | sm100a, sm120a | f0fb1e7e22e6dc0232e42522d01260b5be5dd1d4592b28acb6935431b4e0cac7 | 2,764,183 |
| sageattention-odyssey.patch | source (git am onto d1a57a5, tree c1f23d3) |
dd86a4dfb3689f2cfe7bd4dd6ad4e78c38d1b1ec7f83ce66ed4ac235b0718156 | 21,585 |
Rebuilding
git clone https://github.com/thu-ml/SageAttention.git && cd SageAttention && git checkout d1a57a5
curl -sSfL https://huggingface.co/amarkovich/attention-wheels/resolve/main/wheels/odyssey-acd75b5/sageattention-odyssey.patch | git am
export CUDA_HOME=<cuda-13.0> PATH=<cuda-13.0>/bin:$PATH SAGEATTENTION_LOCAL_VERSION=cu130torch2.9cxx11abiTRUE
TORCH_CUDA_ARCH_LIST="8.0;8.9;9.0;12.0" EXT_PARALLEL=4 MAX_JOBS=8 python setup.py bdist_wheel
(cd sageattention3_blackwell && TORCH_CUDA_ARCH_LIST="10.0;12.0" MAX_JOBS=4 python setup.py bdist_wheel)
Run it with a Python that has torch==2.9.1+cu130, setuptools<75, wheel<0.44, packaging, and ninja installed.
Adding a wheel
Build from pinned source, upload under wheels/<tag>/ with the patch that produced it, verify the sha256 by downloading through the resolve/ URL (an HTTP 200 is not integrity evidence), and add a row above.