pin v2 payload: engine 2610071254001, cuda 2610071451001, cuda12 2610071451002, sbsa 2610071452001, vulkan 2610071443001, macos 2610041000001, media 2610040945001
Browse files- .gitattributes +8 -0
- components/cuda/2610071451001/BUILD_INFO.cuda.md +33 -0
- components/cuda/2610071451001/ggml-cuda-win-x86_64.dll +3 -0
- components/cuda/2610071451001/ggml-cuda-x86_64.so +3 -0
- components/cuda12/2610071451002/BUILD_INFO.cuda12.md +30 -0
- components/cuda12/2610071451002/ggml-cuda-cu12-x86_64.so +3 -0
- components/engine/2610071254001/BUILD_INFO.engine.md +55 -0
- components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001 +3 -0
- components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001.exe +3 -0
- components/sbsa/2610071452001/BUILD_INFO.sbsa.md +28 -0
- components/sbsa/2610071452001/ggml-cuda-sbsa-aarch64.so +3 -0
- components/vulkan/2610071443001/BUILD_INFO.vulkan.md +37 -0
- components/vulkan/2610071443001/ggml-vulkan-win-x86_64.dll +3 -0
- components/vulkan/2610071443001/ggml-vulkan-x86_64.so +3 -0
.gitattributes
CHANGED
|
@@ -236,3 +236,11 @@ components/vulkan/2610041656001/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=
|
|
| 236 |
components/vulkan/2610041656001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
|
| 237 |
components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001 filter=lfs diff=lfs merge=lfs -text
|
| 238 |
components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001.exe filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 236 |
components/vulkan/2610041656001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
|
| 237 |
components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001 filter=lfs diff=lfs merge=lfs -text
|
| 238 |
components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001.exe filter=lfs diff=lfs merge=lfs -text
|
| 239 |
+
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001 filter=lfs diff=lfs merge=lfs -text
|
| 240 |
+
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001.exe filter=lfs diff=lfs merge=lfs -text
|
| 241 |
+
components/cuda/2610071451001/ggml-cuda-x86_64.so filter=lfs diff=lfs merge=lfs -text
|
| 242 |
+
components/cuda/2610071451001/ggml-cuda-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
|
| 243 |
+
components/cuda12/2610071451002/ggml-cuda-cu12-x86_64.so filter=lfs diff=lfs merge=lfs -text
|
| 244 |
+
components/sbsa/2610071452001/ggml-cuda-sbsa-aarch64.so filter=lfs diff=lfs merge=lfs -text
|
| 245 |
+
components/vulkan/2610071443001/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=lfs -text
|
| 246 |
+
components/vulkan/2610071443001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
|
components/cuda/2610071451001/BUILD_INFO.cuda.md
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# cuda 2610071451001
|
| 2 |
+
|
| 3 |
+
Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
|
| 4 |
+
|
| 5 |
+
Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with 2610071254001.
|
| 6 |
+
|
| 7 |
+
| file | platform | kind | sha256 | bytes |
|
| 8 |
+
|---|---|---|---|---|
|
| 9 |
+
| `ggml-cuda-x86_64.so` | x86_64 | cuda | `05d3a1544130704b2c116e6506f3ae65fb3bbd2563e90fe96541854e4c8ba6d5` | 725253760 |
|
| 10 |
+
| `ggml-cuda-win-x86_64.dll` | win-x86_64 | cuda | `654a8f5f20c4f420d3a8f63025b55df71c2f4cb15383a3386203e482a10d17c0` | 700748800 |
|
| 11 |
+
|
| 12 |
+
ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
|
| 13 |
+
|
| 14 |
+
```
|
| 15 |
+
llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
|
| 16 |
+
llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
|
| 17 |
+
llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
|
| 18 |
+
llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
|
| 19 |
+
llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
|
| 20 |
+
llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
|
| 21 |
+
llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
|
| 22 |
+
llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
|
| 23 |
+
llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
|
| 24 |
+
llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
|
| 25 |
+
llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
|
| 26 |
+
llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
|
| 27 |
+
```
|
| 28 |
+
|
| 29 |
+
CUDA 13. Six generations, 402 kernel images each; `120` is the sm_120f kernel set.
|
| 30 |
+
|
| 31 |
+
**The two files are of different builds.** The Linux file is built at patch chain 0585 and reports each device's PCI address to the engine (0584; on Linux the engine still takes the link from nvidia-smi). The Windows file is the six-generation build of the 0583 tree: it predates 0584, so on Windows `OPENCOTI_LINK_GBPS` by PCI address and the `source=system` link profile need the next CUDA library (building now); with this file the engine behaves as before (`OPENCOTI_LINK_GBPS` with a plain number works).
|
| 32 |
+
|
| 33 |
+
Gated: the Linux file on bs2 (sm_120) with engine 2610071254001 — capped and uncapped KVarN (v6, 11.5 GB, 8k / 59k / 88k, one slot and `-kvu`), same peaks and decode as the 0583 release library — and against glibc 2.28. The Windows file: Windows 11, RTX 3090 — KVarN 4-bit one slot, cap 14,500 MiB: capped = uncapped at 4k / 40k / 62k tokens, max log-probability difference 0.000000; peak 13,310 MiB.
|
components/cuda/2610071451001/ggml-cuda-win-x86_64.dll
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:654a8f5f20c4f420d3a8f63025b55df71c2f4cb15383a3386203e482a10d17c0
|
| 3 |
+
size 700748800
|
components/cuda/2610071451001/ggml-cuda-x86_64.so
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:05d3a1544130704b2c116e6506f3ae65fb3bbd2563e90fe96541854e4c8ba6d5
|
| 3 |
+
size 725253760
|
components/cuda12/2610071451002/BUILD_INFO.cuda12.md
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# cuda12 2610071451002
|
| 2 |
+
|
| 3 |
+
Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
|
| 4 |
+
|
| 5 |
+
Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with nothing — NOT run.
|
| 6 |
+
|
| 7 |
+
| file | platform | kind | sha256 | bytes |
|
| 8 |
+
|---|---|---|---|---|
|
| 9 |
+
| `ggml-cuda-cu12-x86_64.so` | x86_64 | cuda12 | `08efa174b7dda7510649f45016b03f72f5ee70ff546554f7ec6dd3aa28b1e250` | 250856912 |
|
| 10 |
+
|
| 11 |
+
ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
|
| 12 |
+
|
| 13 |
+
```
|
| 14 |
+
llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
|
| 15 |
+
llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
|
| 16 |
+
llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
|
| 17 |
+
llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
|
| 18 |
+
llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
|
| 19 |
+
llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
|
| 20 |
+
llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
|
| 21 |
+
llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
|
| 22 |
+
llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
|
| 23 |
+
llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
|
| 24 |
+
llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
|
| 25 |
+
llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
CUDA 12.8, for GPUs older than sm_75 (Maxwell, Pascal, Volta). Built at patch chain 0585. Carries the bug-3932 fix (q4_0 / q8_0 KV on GPUs without tensor cores).
|
| 29 |
+
|
| 30 |
+
**These bytes have not been run on a card.** The bug-3932 fix was run on a V100 with an earlier build of the same lane.
|
components/cuda12/2610071451002/ggml-cuda-cu12-x86_64.so
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:08efa174b7dda7510649f45016b03f72f5ee70ff546554f7ec6dd3aa28b1e250
|
| 3 |
+
size 250856912
|
components/engine/2610071254001/BUILD_INFO.engine.md
ADDED
|
@@ -0,0 +1,55 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# engine 2610071254001
|
| 2 |
+
|
| 3 |
+
Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
|
| 4 |
+
|
| 5 |
+
| file | platform | kind | sha256 | bytes |
|
| 6 |
+
|---|---|---|---|---|
|
| 7 |
+
| `opencoti-0.10.5-c9-2610071254001` | any | bin | `bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b` | 134866740 |
|
| 8 |
+
| `opencoti-0.10.5-c9-2610071254001.exe` | win-x86_64 | bin | `bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b` | 134866740 |
|
| 9 |
+
|
| 10 |
+
ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
|
| 11 |
+
|
| 12 |
+
```
|
| 13 |
+
llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
|
| 14 |
+
llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
|
| 15 |
+
llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
|
| 16 |
+
llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
|
| 17 |
+
llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
|
| 18 |
+
llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
|
| 19 |
+
llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
|
| 20 |
+
llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
|
| 21 |
+
llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
|
| 22 |
+
llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
|
| 23 |
+
llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
|
| 24 |
+
llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
|
| 25 |
+
```
|
| 26 |
+
|
| 27 |
+
ABI `media` = `b099409b3029a6e380268e31d1caefa6706569961d32e72b3543fbc73903d95c` (exported) — sha256 of this list:
|
| 28 |
+
|
| 29 |
+
```
|
| 30 |
+
llamafile/oc-audiocpp/audiocpp.h 6ad0b55d5f28adc9ad60a45bf9a7caa428c4f5441389a6932d6670a8e2a90d73
|
| 31 |
+
llamafile/oc-codec/oc_codec_sidecar.h f2c9001d1414aa2a55533b867b79c52a52c682558e0234c5bebd30bf97b91005
|
| 32 |
+
```
|
| 33 |
+
|
| 34 |
+
ABI `loader` = `d6ff34022f1d3fb5af56ca0c80e3f4e552bf09f1bebdf0cb5f314a454bf38dbc` (exported) — sha256 of this list:
|
| 35 |
+
|
| 36 |
+
```
|
| 37 |
+
.cosmocc/4.0.2/bin/ape-m1.c 78af79e20abbb3f99a97355959b59bd58949564064adb754f0302ebabb4bf7a3
|
| 38 |
+
```
|
| 39 |
+
|
| 40 |
+
The APE: one file for Linux x86_64 / aarch64, Windows and macOS arm64; the `.exe` row is the same bytes under the name Windows needs.
|
| 41 |
+
|
| 42 |
+
**The c10 development line, for integration testing. Not a release.** Patch chain 0001…0584 (the c9 release is 0561). Replaces dev engine 2610041714001.
|
| 43 |
+
|
| 44 |
+
What is new since c9:
|
| 45 |
+
|
| 46 |
+
- **The KV rolling window is ONE sliding window with an adaptive size, and it engages only when the KV cache does not fit VRAM** (0562…0568, 0576…0582). Below the limit a capped boot gives the same output as an uncapped boot, bit for bit: f16, q4_0, q8_0 and KVarN caches, one slot and `-kvu`, CUDA and Vulkan. Past the limit the cache keeps its resident part in VRAM and reads the rest from host memory through the window.
|
| 47 |
+
- **KVarN under a VRAM cap** (bug-3943, 0576…0582): the records stay on the device, the window reads them staged; a context shift on a capped KVarN cache runs on the device (bug-3948, 0582).
|
| 48 |
+
- **Two Vulkan cards under a cap** (bug-3950, 0583): a capped boot on two Vulkan devices aborted at start; fixed. The fix crashed Linux boots with two Vulkan devices (bug-3953); the Vulkan library of this index carries 0585, which fixes that.
|
| 49 |
+
- **`OPENCOTI_LINK_GBPS` and the link profile on Windows** (bug-3952, 0584): `--list-devices` with `OPENCOTI_LIST_DEVICE_IDS=1` prints ` pci=<address>` after `id=`; an `OPENCOTI_LINK_GBPS` key matches a device by its id or by its PCI address; the link profile reads the PCIe link from the system (`source=system`). The device id itself is unchanged. On Windows this needs a GPU library built at 0584 or later: the Vulkan library of this index is; the Windows CUDA library of this index is not (0583 build) and reports no PCI address — the engine then behaves as before. The next Windows CUDA library carries it.
|
| 50 |
+
- **Cloudflare Clef decision head** (0573): `POST /embedding` with `score_fields`; feature `clef_score_v1`.
|
| 51 |
+
- **Media**: flash attention is on by default for image and video engines (0570); a C++ throw inside the Windows Vulkan library no longer kills the engine (bug-3937, 0572 — the RX 9070 XT video crash); Vulkan drivers too old to work are refused with a reason (0571, `OPENCOTI_VK_ALLOW_OLD_DRIVER=1` keeps them).
|
| 52 |
+
- **q4_0 / q8_0 KV on GPUs without tensor cores** (bug-3932, 0565): fixed; needs the legacy CUDA 12 library of this index.
|
| 53 |
+
- A model that cannot be loaded no longer aborts the fit probe (bug-3938, 0574).
|
| 54 |
+
|
| 55 |
+
Gated: bs2 (sm_120) on CUDA and Vulkan against glibc 2.28; Windows 11 on an RTX 3090 (CUDA, with the previous dev CUDA library) and an RX 9070 XT (Vulkan), and both cards together through Vulkan; the capped = uncapped equality runs are in the patches README (0580…0584).
|
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b
|
| 3 |
+
size 134866740
|
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001.exe
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b
|
| 3 |
+
size 134866740
|
components/sbsa/2610071452001/BUILD_INFO.sbsa.md
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# sbsa 2610071452001
|
| 2 |
+
|
| 3 |
+
Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
|
| 4 |
+
|
| 5 |
+
Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with nothing — NOT run.
|
| 6 |
+
|
| 7 |
+
| file | platform | kind | sha256 | bytes |
|
| 8 |
+
|---|---|---|---|---|
|
| 9 |
+
| `ggml-cuda-sbsa-aarch64.so` | aarch64 | cuda | `54378184340b2ef5d1680c26d88fa89813acea2f127b18bbdec4276958bea4f7` | 333095296 |
|
| 10 |
+
|
| 11 |
+
ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
|
| 12 |
+
|
| 13 |
+
```
|
| 14 |
+
llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
|
| 15 |
+
llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
|
| 16 |
+
llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
|
| 17 |
+
llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
|
| 18 |
+
llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
|
| 19 |
+
llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
|
| 20 |
+
llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
|
| 21 |
+
llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
|
| 22 |
+
llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
|
| 23 |
+
llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
|
| 24 |
+
llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
|
| 25 |
+
llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
CUDA 13 for Linux aarch64 (SBSA). Built at patch chain 0585. **Never run: there is no such machine in the test fleet.** Reports from users are the gate.
|
components/sbsa/2610071452001/ggml-cuda-sbsa-aarch64.so
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:54378184340b2ef5d1680c26d88fa89813acea2f127b18bbdec4276958bea4f7
|
| 3 |
+
size 333095296
|
components/vulkan/2610071443001/BUILD_INFO.vulkan.md
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# vulkan 2610071443001
|
| 2 |
+
|
| 3 |
+
Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
|
| 4 |
+
|
| 5 |
+
Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with 2610071254001.
|
| 6 |
+
|
| 7 |
+
| file | platform | kind | sha256 | bytes |
|
| 8 |
+
|---|---|---|---|---|
|
| 9 |
+
| `ggml-vulkan-x86_64.so` | x86_64 | vulkan | `b34a07cb706706ecc44c50bdeff0ac5d03545ea78b5c0da9b00acc7584d16281` | 56322272 |
|
| 10 |
+
| `ggml-vulkan-win-x86_64.dll` | win-x86_64 | vulkan | `ea7a1c51d9e8922cdff73355fb734859a168b5c621299acb5332d21c9ea02e76` | 57352976 |
|
| 11 |
+
|
| 12 |
+
ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
|
| 13 |
+
|
| 14 |
+
```
|
| 15 |
+
llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
|
| 16 |
+
llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
|
| 17 |
+
llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
|
| 18 |
+
llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
|
| 19 |
+
llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
|
| 20 |
+
llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
|
| 21 |
+
llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
|
| 22 |
+
llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
|
| 23 |
+
llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
|
| 24 |
+
llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
|
| 25 |
+
llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
|
| 26 |
+
llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
|
| 27 |
+
```
|
| 28 |
+
|
| 29 |
+
Built at patch chain 0585 (one patch past the engine of this index: 0585 changes this library only). One host buffer type per device (bug-3950, 0583): a capped boot on two Vulkan cards starts. **0585 (bug-3953)**: the 0583 library crashed every LINUX boot that saw two Vulkan devices (exit 139 at model load); fixed here, never published before the fix. The KVarN record window on Vulkan (0577…0581). A C++ throw inside the Windows library is dispatched by the library itself (bug-3937, 0572). Drivers too old to work are refused with a reason (0571). On Windows the library reports each device's PCI address and PCIe link to the engine (0584).
|
| 30 |
+
|
| 31 |
+
A Vulkan device's id is unchanged: its PCI address, or `uuid:<deviceUUID>` where the driver has no `VK_EXT_pci_bus_info` (Windows).
|
| 32 |
+
|
| 33 |
+
Gated: bs2 (sm_120, NVIDIA) against glibc 2.28 (old-glibc gate PASS), and on two Vulkan devices on Linux (RTX PRO 6000 + llvmpipe: boots, answers, exits 0). Windows 11: the 0583/0584 libraries on an RX 9070 XT, an RTX 3090 and both cards together; this 0585 Windows file on both cards: KVarN one slot / 4 slots and q4_0 give the same results as the 0583 file to six digits, control aborts as before.
|
| 34 |
+
|
| 35 |
+
AMD Radeon on Windows: AMD Software 26.9.
|
| 36 |
+
|
| 37 |
+
**Known open (bug-3954):** two Vulkan cards, q4_0 KV, a capped boot vs an uncapped boot with the SAME layer split (`-ts 1,1`): the same text, but the log-probabilities differ by up to 0.21 at 4k tokens — below the limit, where they must be equal. KVarN on two cards is exact there; q4_0 on one card is exact (bs2, Vulkan and CUDA). Being isolated. Without a fixed split, a per-card cap also moves layers between the cards, so capped and uncapped boots never compute bit-identically on two cards — that is placement, not a defect.
|
components/vulkan/2610071443001/ggml-vulkan-win-x86_64.dll
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ea7a1c51d9e8922cdff73355fb734859a168b5c621299acb5332d21c9ea02e76
|
| 3 |
+
size 57352976
|
components/vulkan/2610071443001/ggml-vulkan-x86_64.so
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b34a07cb706706ecc44c50bdeff0ac5d03545ea78b5c0da9b00acc7584d16281
|
| 3 |
+
size 56322272
|