|
Download README2.md from rththr/llamacpp-binary: direct link, hf CLI and curl.
- Browser
- Download file 1.32 kB
-
https://huggingface.co/rththr/llamacpp-binary/resolve/main/README2.md
- Command line
-
hf download hf://rththr/llamacpp-binary/README2.md
-
curl -L -o README2.md https://huggingface.co/rththr/llamacpp-binary/resolve/main/README2.md
1.32 kB
| # SageAttention 2.2.0 — Pre-built Wheel | |
| Pre-compiled SageAttention 2.2.0 wheel for **Linux x86_64** with CUDA 13 and C++20. | |
| ## Compatibility | |
| | Component | Version | | |
| |-----------|---------| | |
| | Python | 3.12 | | |
| | PyTorch | 2.15.0 nightly (cu132) | | |
| | CUDA Toolkit | 13.3 | | |
| | GPU | Ada Lovelace (sm_89) — RTX 4050/4060/4070/4080/4090 | | |
| | SageAttention | 2.2.0 | | |
| ## Install | |
| ```bash | |
| pip install sageattention-2.2.0-cp312-cp312-linux_x86_64.whl --no-deps | |
| ``` | |
| ## Build from Source (if needed) | |
| ### Requirements | |
| - CUDA Toolkit (nvcc) | |
| - PyTorch with CUDA support | |
| - 16GB+ RAM or adequate swap | |
| ### Steps | |
| ```bash | |
| git clone --depth 1 https://github.com/thu-ml/SageAttention.git /tmp/SageAttention | |
| cd /tmp/SageAttention | |
| # Patch: PyTorch 2.15 nightly requires C++20 | |
| sed -i 's/-std=c++17/-std=c++20/g' setup.py | |
| # Build | |
| CUDA_HOME=/usr/local/cuda-13.3 \ | |
| TORCH_CUDA_ARCH_LIST="8.9" \ | |
| MAX_JOBS=1 \ | |
| NVCC_APPEND_FLAGS="--threads 8" \ | |
| python setup.py install | |
| ``` | |
| ### Create Wheel | |
| ```bash | |
| python setup.py bdist_wheel | |
| cp dist/*.whl /path/to/wheels/ | |
| ``` | |
| ## Notes | |
| - Built on Ubuntu 24.04 with 14GB RAM + 19GB swap | |
| - `MAX_JOBS=1` used to avoid OOM during CUDA kernel compilation | |
| - PyTorch 2.15 nightly requires C++20; upstream SageAttention ships with C++17 which must be patched | |
| - Wheel includes sm_80 (Ampere) and sm_89 (Ada) kernels | |