Instructions to use replicate/sage-attention with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/sage-attention with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("replicate/sage-attention", version=1) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from replicate/sage-attention: direct link, hf CLI and curl.
- Browser
- Download file 1.1 kB
-
https://huggingface.co/replicate/sage-attention/resolve/main/README.md
- Command line
-
hf download hf://replicate/sage-attention/README.md
-
curl -L -o README.md https://huggingface.co/replicate/sage-attention/resolve/main/README.md
1.1 kB
| library_name: kernels | |
| license: apache-2.0 | |
| > [!CAUTION] | |
| > Starting from September 13, 2026, we will be removing the "model" type repositories of kernels (e.g., kernels-community/flash-attn3). Make sure you're using a latest version of kernels. If you face any disruption, please report them here: https://github.com/huggingface/kernels/issues/new. | |
| This is the repository card of kernels-community/sage-attention that has been pushed on the Hub. It was built to be used with the [`kernels` library](https://github.com/huggingface/kernels). This card was automatically generated. | |
| ## How to use | |
| ```python | |
| # make sure `kernels` is installed: `pip install -U kernels` | |
| from kernels import get_kernel | |
| kernel_module = get_kernel("kernels-community/sage-attention") | |
| per_block_int8 = kernel_module.per_block_int8 | |
| per_block_int8(...) | |
| ``` | |
| ## Available functions | |
| - `per_block_int8` | |
| - `per_warp_int8` | |
| - `sub_mean` | |
| - `per_channel_fp8` | |
| - `sageattn` | |
| - `sageattn3_blackwell` | |
| ## Benchmarks | |
| Benchmarking script is available for this kernel. Run `kernels benchmark kernels-community/sage-attention`. | |