|
Download README2.md from rththr/llamacpp-binary: direct link, hf CLI and curl.
- Browser
- Download file 1.32 kB
-
https://huggingface.co/rththr/llamacpp-binary/resolve/main/README2.md
- Command line
-
hf download hf://rththr/llamacpp-binary/README2.md
-
curl -L -o README2.md https://huggingface.co/rththr/llamacpp-binary/resolve/main/README2.md
1.32 kB
SageAttention 2.2.0 — Pre-built Wheel
Pre-compiled SageAttention 2.2.0 wheel for Linux x86_64 with CUDA 13 and C++20.
Compatibility
| Component | Version |
|---|---|
| Python | 3.12 |
| PyTorch | 2.15.0 nightly (cu132) |
| CUDA Toolkit | 13.3 |
| GPU | Ada Lovelace (sm_89) — RTX 4050/4060/4070/4080/4090 |
| SageAttention | 2.2.0 |
Install
pip install sageattention-2.2.0-cp312-cp312-linux_x86_64.whl --no-deps
Build from Source (if needed)
Requirements
- CUDA Toolkit (nvcc)
- PyTorch with CUDA support
- 16GB+ RAM or adequate swap
Steps
git clone --depth 1 https://github.com/thu-ml/SageAttention.git /tmp/SageAttention
cd /tmp/SageAttention
# Patch: PyTorch 2.15 nightly requires C++20
sed -i 's/-std=c++17/-std=c++20/g' setup.py
# Build
CUDA_HOME=/usr/local/cuda-13.3 \
TORCH_CUDA_ARCH_LIST="8.9" \
MAX_JOBS=1 \
NVCC_APPEND_FLAGS="--threads 8" \
python setup.py install
Create Wheel
python setup.py bdist_wheel
cp dist/*.whl /path/to/wheels/
Notes
- Built on Ubuntu 24.04 with 14GB RAM + 19GB swap
MAX_JOBS=1used to avoid OOM during CUDA kernel compilation- PyTorch 2.15 nightly requires C++20; upstream SageAttention ships with C++17 which must be patched
- Wheel includes sm_80 (Ampere) and sm_89 (Ada) kernels