How to use from the
Use from the
Kernels library
# !pip install kernels

from kernels import get_kernel

# a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones
kernel = get_kernel("pat883/attn-decode", version=1)

Flash-Attention Decode (RDNA4/gfx1201)

Dense, paged, and fp8 flash-attention decode kernels (GQA, sliding-window, block-table KV cache).

Built with kernel-builder for AMD RDNA4 (gfx1201). Load with the kernels library:

from kernels import get_kernel
kernel = get_kernel("pat883/attn-decode")

Requires a ROCm PyTorch build (torch 2.10 / ROCm 7.x) on an RDNA4 card. Built variants: torch210-cxx11-rocm70 and torch210-cxx11-rocm71.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support