YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SageAttention 2.2.0 β RTX 50 Series Blackwell (Linux)
Pre-built SageAttention 2.2.0 for NVIDIA RTX 5090 (sm_120 Blackwell) on Linux. INT8 quantized attention achieving 2-3x speedup over FlashAttention2 without accuracy loss.
Built from woct0rdho/SageAttention β a fork of thu-ml/SageAttention with ABI3 stable wheel support and additional fixes.
Wheel
sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl
Requirements
| Component | Version |
|---|---|
| GPU | NVIDIA RTX 50 series (sm_120 Blackwell) |
| OS | Linux x86_64 |
| Python | 3.9+ (ABI3 stable β compatible with Python 3.9 or higher) |
| PyTorch | 2.11.0 or higher |
| CUDA | 13.0 or higher |
| Triton | compatible with your PyTorch version |
Installation
pip install sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl --no-deps
Important: Always use
--no-depsto prevent pip from overwriting your PyTorch installation.
Verify
import torch
print(torch.__version__)
from sageattention import sageattn_varlen
print("SageAttention 2.2.0 OK")
Features
- sageattn_varlen: Variable-length attention (used by SeedVR2, WanVideo, etc.)
- INT8 QK quantization with per-thread granularity
- FP16 PV accumulation for accuracy
- Supports sm_80 (Ampere), sm_89 (Ada), sm_120 (Blackwell) via included CUDA kernels + Triton fallback
- Native varlen API β no reshaping overhead
- ABI3 stable wheel β single wheel compatible with Python 3.9 and above
Source
- Fork: woct0rdho/SageAttention
- Upstream: thu-ml/SageAttention v2.2.0
License
Apache 2.0
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support