Instructions to use replicate/paged-attention with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/paged-attention with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("replicate/paged-attention", version=1) - Notebooks
- Google Colab
- Kaggle
Download cuda-utils/cuda_utils.h from replicate/paged-attention: direct link, hf CLI and curl.
- Browser
- Download file 1.41 kB
-
https://huggingface.co/replicate/paged-attention/resolve/main/cuda-utils/cuda_utils.h
- Command line
-
hf download hf://replicate/paged-attention/cuda-utils/cuda_utils.h
-
curl -L -o cuda_utils.h https://huggingface.co/replicate/paged-attention/resolve/main/cuda-utils/cuda_utils.h
1.41 kB
| int64_t get_device_attribute(int64_t attribute, int64_t device_id); | |
| int64_t get_max_shared_memory_per_block_device_attribute(int64_t device_id); | |
| namespace cuda_utils { | |
| template <typename T> | |
| HOST_DEVICE_INLINE constexpr std::enable_if_t<std::is_integral_v<T>, T> | |
| ceil_div(T a, T b) { | |
| return (a + b - 1) / b; | |
| } | |
| }; // namespace cuda_utils |