ManniX-ITA commited on
Commit
a60ac89
·
verified ·
1 Parent(s): ad2b7db

pin v2 payload: engine 2610071254001, cuda 2610071451001, cuda12 2610071451002, sbsa 2610071452001, vulkan 2610071443001, macos 2610041000001, media 2610040945001

Browse files
.gitattributes CHANGED
@@ -236,3 +236,11 @@ components/vulkan/2610041656001/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=
236
  components/vulkan/2610041656001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
237
  components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001 filter=lfs diff=lfs merge=lfs -text
238
  components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001.exe filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
236
  components/vulkan/2610041656001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
237
  components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001 filter=lfs diff=lfs merge=lfs -text
238
  components/engine/2610041714001/opencoti-0.10.5-c7-2610041714001.exe filter=lfs diff=lfs merge=lfs -text
239
+ components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001 filter=lfs diff=lfs merge=lfs -text
240
+ components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001.exe filter=lfs diff=lfs merge=lfs -text
241
+ components/cuda/2610071451001/ggml-cuda-x86_64.so filter=lfs diff=lfs merge=lfs -text
242
+ components/cuda/2610071451001/ggml-cuda-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
243
+ components/cuda12/2610071451002/ggml-cuda-cu12-x86_64.so filter=lfs diff=lfs merge=lfs -text
244
+ components/sbsa/2610071452001/ggml-cuda-sbsa-aarch64.so filter=lfs diff=lfs merge=lfs -text
245
+ components/vulkan/2610071443001/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=lfs -text
246
+ components/vulkan/2610071443001/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
components/cuda/2610071451001/BUILD_INFO.cuda.md ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # cuda 2610071451001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with 2610071254001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-x86_64.so` | x86_64 | cuda | `05d3a1544130704b2c116e6506f3ae65fb3bbd2563e90fe96541854e4c8ba6d5` | 725253760 |
10
+ | `ggml-cuda-win-x86_64.dll` | win-x86_64 | cuda | `654a8f5f20c4f420d3a8f63025b55df71c2f4cb15383a3386203e482a10d17c0` | 700748800 |
11
+
12
+ ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
13
+
14
+ ```
15
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
16
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
17
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
18
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
19
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
20
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
21
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
22
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
23
+ llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
24
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
25
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
26
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
27
+ ```
28
+
29
+ CUDA 13. Six generations, 402 kernel images each; `120` is the sm_120f kernel set.
30
+
31
+ **The two files are of different builds.** The Linux file is built at patch chain 0585 and reports each device's PCI address to the engine (0584; on Linux the engine still takes the link from nvidia-smi). The Windows file is the six-generation build of the 0583 tree: it predates 0584, so on Windows `OPENCOTI_LINK_GBPS` by PCI address and the `source=system` link profile need the next CUDA library (building now); with this file the engine behaves as before (`OPENCOTI_LINK_GBPS` with a plain number works).
32
+
33
+ Gated: the Linux file on bs2 (sm_120) with engine 2610071254001 — capped and uncapped KVarN (v6, 11.5 GB, 8k / 59k / 88k, one slot and `-kvu`), same peaks and decode as the 0583 release library — and against glibc 2.28. The Windows file: Windows 11, RTX 3090 — KVarN 4-bit one slot, cap 14,500 MiB: capped = uncapped at 4k / 40k / 62k tokens, max log-probability difference 0.000000; peak 13,310 MiB.
components/cuda/2610071451001/ggml-cuda-win-x86_64.dll ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:654a8f5f20c4f420d3a8f63025b55df71c2f4cb15383a3386203e482a10d17c0
3
+ size 700748800
components/cuda/2610071451001/ggml-cuda-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:05d3a1544130704b2c116e6506f3ae65fb3bbd2563e90fe96541854e4c8ba6d5
3
+ size 725253760
components/cuda12/2610071451002/BUILD_INFO.cuda12.md ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # cuda12 2610071451002
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with nothing — NOT run.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-cu12-x86_64.so` | x86_64 | cuda12 | `08efa174b7dda7510649f45016b03f72f5ee70ff546554f7ec6dd3aa28b1e250` | 250856912 |
10
+
11
+ ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
12
+
13
+ ```
14
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
15
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
16
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
17
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
18
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
19
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
20
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
21
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
22
+ llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
23
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
24
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
25
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
26
+ ```
27
+
28
+ CUDA 12.8, for GPUs older than sm_75 (Maxwell, Pascal, Volta). Built at patch chain 0585. Carries the bug-3932 fix (q4_0 / q8_0 KV on GPUs without tensor cores).
29
+
30
+ **These bytes have not been run on a card.** The bug-3932 fix was run on a V100 with an earlier build of the same lane.
components/cuda12/2610071451002/ggml-cuda-cu12-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:08efa174b7dda7510649f45016b03f72f5ee70ff546554f7ec6dd3aa28b1e250
3
+ size 250856912
components/engine/2610071254001/BUILD_INFO.engine.md ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # engine 2610071254001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ | file | platform | kind | sha256 | bytes |
6
+ |---|---|---|---|---|
7
+ | `opencoti-0.10.5-c9-2610071254001` | any | bin | `bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b` | 134866740 |
8
+ | `opencoti-0.10.5-c9-2610071254001.exe` | win-x86_64 | bin | `bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b` | 134866740 |
9
+
10
+ ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
11
+
12
+ ```
13
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
14
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
15
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
16
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
17
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
18
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
19
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
20
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
21
+ llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
22
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
23
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
24
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
25
+ ```
26
+
27
+ ABI `media` = `b099409b3029a6e380268e31d1caefa6706569961d32e72b3543fbc73903d95c` (exported) — sha256 of this list:
28
+
29
+ ```
30
+ llamafile/oc-audiocpp/audiocpp.h 6ad0b55d5f28adc9ad60a45bf9a7caa428c4f5441389a6932d6670a8e2a90d73
31
+ llamafile/oc-codec/oc_codec_sidecar.h f2c9001d1414aa2a55533b867b79c52a52c682558e0234c5bebd30bf97b91005
32
+ ```
33
+
34
+ ABI `loader` = `d6ff34022f1d3fb5af56ca0c80e3f4e552bf09f1bebdf0cb5f314a454bf38dbc` (exported) — sha256 of this list:
35
+
36
+ ```
37
+ .cosmocc/4.0.2/bin/ape-m1.c 78af79e20abbb3f99a97355959b59bd58949564064adb754f0302ebabb4bf7a3
38
+ ```
39
+
40
+ The APE: one file for Linux x86_64 / aarch64, Windows and macOS arm64; the `.exe` row is the same bytes under the name Windows needs.
41
+
42
+ **The c10 development line, for integration testing. Not a release.** Patch chain 0001…0584 (the c9 release is 0561). Replaces dev engine 2610041714001.
43
+
44
+ What is new since c9:
45
+
46
+ - **The KV rolling window is ONE sliding window with an adaptive size, and it engages only when the KV cache does not fit VRAM** (0562…0568, 0576…0582). Below the limit a capped boot gives the same output as an uncapped boot, bit for bit: f16, q4_0, q8_0 and KVarN caches, one slot and `-kvu`, CUDA and Vulkan. Past the limit the cache keeps its resident part in VRAM and reads the rest from host memory through the window.
47
+ - **KVarN under a VRAM cap** (bug-3943, 0576…0582): the records stay on the device, the window reads them staged; a context shift on a capped KVarN cache runs on the device (bug-3948, 0582).
48
+ - **Two Vulkan cards under a cap** (bug-3950, 0583): a capped boot on two Vulkan devices aborted at start; fixed. The fix crashed Linux boots with two Vulkan devices (bug-3953); the Vulkan library of this index carries 0585, which fixes that.
49
+ - **`OPENCOTI_LINK_GBPS` and the link profile on Windows** (bug-3952, 0584): `--list-devices` with `OPENCOTI_LIST_DEVICE_IDS=1` prints ` pci=<address>` after `id=`; an `OPENCOTI_LINK_GBPS` key matches a device by its id or by its PCI address; the link profile reads the PCIe link from the system (`source=system`). The device id itself is unchanged. On Windows this needs a GPU library built at 0584 or later: the Vulkan library of this index is; the Windows CUDA library of this index is not (0583 build) and reports no PCI address — the engine then behaves as before. The next Windows CUDA library carries it.
50
+ - **Cloudflare Clef decision head** (0573): `POST /embedding` with `score_fields`; feature `clef_score_v1`.
51
+ - **Media**: flash attention is on by default for image and video engines (0570); a C++ throw inside the Windows Vulkan library no longer kills the engine (bug-3937, 0572 — the RX 9070 XT video crash); Vulkan drivers too old to work are refused with a reason (0571, `OPENCOTI_VK_ALLOW_OLD_DRIVER=1` keeps them).
52
+ - **q4_0 / q8_0 KV on GPUs without tensor cores** (bug-3932, 0565): fixed; needs the legacy CUDA 12 library of this index.
53
+ - A model that cannot be loaded no longer aborts the fit probe (bug-3938, 0574).
54
+
55
+ Gated: bs2 (sm_120) on CUDA and Vulkan against glibc 2.28; Windows 11 on an RTX 3090 (CUDA, with the previous dev CUDA library) and an RX 9070 XT (Vulkan), and both cards together through Vulkan; the capped = uncapped equality runs are in the patches README (0580…0584).
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b
3
+ size 134866740
components/engine/2610071254001/opencoti-0.10.5-c9-2610071254001.exe ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb4fac5c05dbf69ec29e589e0e79e568b87936a8941414df8761e9700815b44b
3
+ size 134866740
components/sbsa/2610071452001/BUILD_INFO.sbsa.md ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # sbsa 2610071452001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with nothing — NOT run.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-sbsa-aarch64.so` | aarch64 | cuda | `54378184340b2ef5d1680c26d88fa89813acea2f127b18bbdec4276958bea4f7` | 333095296 |
10
+
11
+ ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
12
+
13
+ ```
14
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
15
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
16
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
17
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
18
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
19
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
20
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
21
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
22
+ llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
23
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
24
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
25
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
26
+ ```
27
+
28
+ CUDA 13 for Linux aarch64 (SBSA). Built at patch chain 0585. **Never run: there is no such machine in the test fleet.** Reports from users are the gate.
components/sbsa/2610071452001/ggml-cuda-sbsa-aarch64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:54378184340b2ef5d1680c26d88fa89813acea2f127b18bbdec4276958bea4f7
3
+ size 333095296
components/vulkan/2610071443001/BUILD_INFO.vulkan.md ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # vulkan 2610071443001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610071429001; gated floor (engine-min) 2610071254001; gated with 2610071254001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-vulkan-x86_64.so` | x86_64 | vulkan | `b34a07cb706706ecc44c50bdeff0ac5d03545ea78b5c0da9b00acc7584d16281` | 56322272 |
10
+ | `ggml-vulkan-win-x86_64.dll` | win-x86_64 | vulkan | `ea7a1c51d9e8922cdff73355fb734859a168b5c621299acb5332d21c9ea02e76` | 57352976 |
11
+
12
+ ABI `ggml` = `41ba479c24d7e3a624c7b13aa3cbe357e41fac04a504c54ec2ee170de0bff9a1` (exported) — sha256 of this list:
13
+
14
+ ```
15
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
16
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
17
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
18
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
19
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
20
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
21
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
22
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
23
+ llama.cpp/ggml/include/ggml.h f2bb8b127e3e601a78fe32242ffab7235e6a4db3bab5370b72fb6bacba2b2433
24
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
25
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
26
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
27
+ ```
28
+
29
+ Built at patch chain 0585 (one patch past the engine of this index: 0585 changes this library only). One host buffer type per device (bug-3950, 0583): a capped boot on two Vulkan cards starts. **0585 (bug-3953)**: the 0583 library crashed every LINUX boot that saw two Vulkan devices (exit 139 at model load); fixed here, never published before the fix. The KVarN record window on Vulkan (0577…0581). A C++ throw inside the Windows library is dispatched by the library itself (bug-3937, 0572). Drivers too old to work are refused with a reason (0571). On Windows the library reports each device's PCI address and PCIe link to the engine (0584).
30
+
31
+ A Vulkan device's id is unchanged: its PCI address, or `uuid:<deviceUUID>` where the driver has no `VK_EXT_pci_bus_info` (Windows).
32
+
33
+ Gated: bs2 (sm_120, NVIDIA) against glibc 2.28 (old-glibc gate PASS), and on two Vulkan devices on Linux (RTX PRO 6000 + llvmpipe: boots, answers, exits 0). Windows 11: the 0583/0584 libraries on an RX 9070 XT, an RTX 3090 and both cards together; this 0585 Windows file on both cards: KVarN one slot / 4 slots and q4_0 give the same results as the 0583 file to six digits, control aborts as before.
34
+
35
+ AMD Radeon on Windows: AMD Software 26.9.
36
+
37
+ **Known open (bug-3954):** two Vulkan cards, q4_0 KV, a capped boot vs an uncapped boot with the SAME layer split (`-ts 1,1`): the same text, but the log-probabilities differ by up to 0.21 at 4k tokens — below the limit, where they must be equal. KVarN on two cards is exact there; q4_0 on one card is exact (bs2, Vulkan and CUDA). Being isolated. Without a fixed split, a per-card cap also moves layers between the cards, so capped and uncapped boots never compute bit-identically on two cards — that is placement, not a defect.
components/vulkan/2610071443001/ggml-vulkan-win-x86_64.dll ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ea7a1c51d9e8922cdff73355fb734859a168b5c621299acb5332d21c9ea02e76
3
+ size 57352976
components/vulkan/2610071443001/ggml-vulkan-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b34a07cb706706ecc44c50bdeff0ac5d03545ea78b5c0da9b00acc7584d16281
3
+ size 56322272