YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SpargeAttn 0.1.0 β RTX 50 Series Blackwell (Linux)
Pre-built SpargeAttn (Sparse SageAttention2) for NVIDIA RTX 5090 (sm_120 Blackwell) on Linux. Block-sparse attention built on top of SageAttention2 β accelerates models without training.
Built from woct0rdho/SpargeAttn β a fork of thu-ml/SageAttention with ABI3 stable wheel support and additional fixes.
Wheel
spas_sage_attn-0.1.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl
Requirements
| Component | Version |
|---|---|
| GPU | NVIDIA RTX 50 series (sm_120 Blackwell) |
| OS | Linux x86_64 |
| Python | 3.9+ (ABI3 stable β compatible with Python 3.9 or higher) |
| PyTorch | 2.11.0 or higher |
| CUDA | 13.0 or higher |
| Triton | compatible with your PyTorch version |
| SageAttention | 2.2.0 (install SA2 wheel first) |
Installation
# Install SA2 first (dependency)
pip install sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl --no-deps
# Install SpargeAttn
pip install spas_sage_attn-0.1.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl --no-deps
Important: Always use
--no-depsto prevent pip from overwriting your PyTorch installation.
Verify
from spas_sage_attn import spas_sage2_attn_meansim_cuda
print("SpargeAttn OK")
Usage
from spas_sage_attn import spas_sage2_attn_meansim_cuda
# q, k, v: (batch, heads, seq_len, head_dim) in fp16/bf16
output = spas_sage2_attn_meansim_cuda(
q, k, v,
is_causal=False,
smooth_k=True,
tensor_layout="HND",
output_dtype=q.dtype
)
Performance Note
Benchmarks on SeedVR2 show pure SageAttention 2 is faster than SpargeAttn on this workload. SpargeAttn adds sparse pattern computation overhead that exceeds savings for SeedVR2's window attention pattern. SpargeAttn may provide better results on workloads with naturally sparse attention patterns.
Source
- Fork: woct0rdho/SpargeAttn
- Upstream: thu-ml/SageAttention β SpargeAttn module (ICML 2025)
License
Apache 2.0