YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

SageAttention 2.2.0 β€” RTX 50 Series Blackwell (Linux)

Pre-built SageAttention 2.2.0 for NVIDIA RTX 5090 (sm_120 Blackwell) on Linux. INT8 quantized attention achieving 2-3x speedup over FlashAttention2 without accuracy loss.

Built from woct0rdho/SageAttention β€” a fork of thu-ml/SageAttention with ABI3 stable wheel support and additional fixes.

Wheel

sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl

Requirements

Component Version
GPU NVIDIA RTX 50 series (sm_120 Blackwell)
OS Linux x86_64
Python 3.9+ (ABI3 stable β€” compatible with Python 3.9 or higher)
PyTorch 2.11.0 or higher
CUDA 13.0 or higher
Triton compatible with your PyTorch version

Installation

pip install sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl --no-deps

Important: Always use --no-deps to prevent pip from overwriting your PyTorch installation.

Verify

import torch
print(torch.__version__)

from sageattention import sageattn_varlen
print("SageAttention 2.2.0 OK")

Features

  • sageattn_varlen: Variable-length attention (used by SeedVR2, WanVideo, etc.)
  • INT8 QK quantization with per-thread granularity
  • FP16 PV accumulation for accuracy
  • Supports sm_80 (Ampere), sm_89 (Ada), sm_120 (Blackwell) via included CUDA kernels + Triton fallback
  • Native varlen API β€” no reshaping overhead
  • ABI3 stable wheel β€” single wheel compatible with Python 3.9 and above

Source

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support