rththr commited on
Commit
a4c799c
·
verified ·
1 Parent(s): 1b4ed93

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +20 -46
README.md CHANGED
@@ -1,57 +1,31 @@
1
- # SageAttention 2.2.0 — Pre-built Wheel
2
 
3
- Pre-compiled SageAttention 2.2.0 wheel for **Linux x86_64** with CUDA 13 and C++20.
4
 
5
- ## Compatibility
6
 
7
- | Component | Version |
8
- |-----------|---------|
9
- | Python | 3.12 |
10
- | PyTorch | 2.15.0 nightly (cu132) |
11
- | CUDA Toolkit | 13.3 |
12
- | GPU | Ada Lovelace (sm_89) — RTX 4050/4060/4070/4080/4090 |
13
- | SageAttention | 2.2.0 |
14
 
15
- ## Install
16
-
17
- ```bash
18
- pip install sageattention-2.2.0-cp312-cp312-linux_x86_64.whl --no-deps
19
- ```
20
-
21
- ## Build from Source (if needed)
22
 
23
- ### Requirements
24
- - CUDA Toolkit (nvcc)
25
- - PyTorch with CUDA support
26
- - 16GB+ RAM or adequate swap
27
 
28
- ### Steps
29
 
30
- ```bash
31
- git clone --depth 1 https://github.com/thu-ml/SageAttention.git /tmp/SageAttention
32
- cd /tmp/SageAttention
33
-
34
- # Patch: PyTorch 2.15 nightly requires C++20
35
- sed -i 's/-std=c++17/-std=c++20/g' setup.py
36
-
37
- # Build
38
- CUDA_HOME=/usr/local/cuda-13.3 \
39
- TORCH_CUDA_ARCH_LIST="8.9" \
40
- MAX_JOBS=1 \
41
- NVCC_APPEND_FLAGS="--threads 8" \
42
- python setup.py install
43
- ```
44
 
45
- ### Create Wheel
46
 
47
  ```bash
48
- python setup.py bdist_wheel
49
- cp dist/*.whl /path/to/wheels/
50
  ```
51
-
52
- ## Notes
53
-
54
- - Built on Ubuntu 24.04 with 14GB RAM + 19GB swap
55
- - `MAX_JOBS=1` used to avoid OOM during CUDA kernel compilation
56
- - PyTorch 2.15 nightly requires C++20; upstream SageAttention ships with C++17 which must be patched
57
- - Wheel includes sm_80 (Ampere) and sm_89 (Ada) kernels
 
1
+ # K2Horizon llama.cpp CUDA Binary
2
 
3
+ Pre-built llama.cpp binary with CUDA support for Linux x86_64.
4
 
5
+ ## Repository
6
 
7
+ - Source: https://github.com/MBZUAI-IFM/llama.cpp
8
+ - Branch: model/K2Horizon
9
+ - Version: 0.3.0-dev
10
+ - Build: 10671
11
+ - Commit: 35999d101
 
 
12
 
13
+ ## Requirements
 
 
 
 
 
 
14
 
15
+ - Linux x86_64
16
+ - NVIDIA GPU with CUDA support
17
+ - CUDA 12.8+
 
18
 
19
+ ## Files
20
 
21
+ | File | Description |
22
+ |------|-------------|
23
+ | `llama-k2horizon-cuda-linux-x64.tar.gz` | K2Horizon llama.cpp CUDA binary |
24
+ | `llama.cpp-b9628-cuda-12.8-amd64.tar.gz` | llama.cpp build b9628 for CUDA 12.8 |
25
+ | `sageattention-2.2.0-cp312-cp312-linux_x86_64.whl` | SageAttention 2.2.0 pre-built wheel |
 
 
 
 
 
 
 
 
 
26
 
27
+ ## Install
28
 
29
  ```bash
30
+ tar xzf llama-k2horizon-cuda-linux-x64.tar.gz
 
31
  ```