ManniX-ITA commited on
Commit
e850481
·
verified ·
1 Parent(s): 100899a

pin v2 payload: engine 2610100921001, cuda 2610100758001, cuda12 2610100758002, sbsa 2610100758003, vulkan 2610100032004, macos 2610100032005, media 2610071902001

Browse files
.gitattributes CHANGED
@@ -263,3 +263,12 @@ components/media/2610071902001/oc-espeak-linux-x86_64.so filter=lfs diff=lfs mer
263
  components/media/2610071902001/oc-espeak-linux-aarch64.so filter=lfs diff=lfs merge=lfs -text
264
  components/media/2610071902001/oc-espeak-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
265
  components/media/2610071902001/oc-espeak-macos-aarch64.dylib filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
263
  components/media/2610071902001/oc-espeak-linux-aarch64.so filter=lfs diff=lfs merge=lfs -text
264
  components/media/2610071902001/oc-espeak-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
265
  components/media/2610071902001/oc-espeak-macos-aarch64.dylib filter=lfs diff=lfs merge=lfs -text
266
+ components/engine/2610100921001/opencoti-0.10.5-c11-2610100921001 filter=lfs diff=lfs merge=lfs -text
267
+ components/engine/2610100921001/opencoti-0.10.5-c11-2610100921001.exe filter=lfs diff=lfs merge=lfs -text
268
+ components/cuda/2610100758001/ggml-cuda-x86_64.so filter=lfs diff=lfs merge=lfs -text
269
+ components/cuda/2610100758001/ggml-cuda-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
270
+ components/cuda12/2610100758002/ggml-cuda-cu12-x86_64.so filter=lfs diff=lfs merge=lfs -text
271
+ components/sbsa/2610100758003/ggml-cuda-sbsa-aarch64.so filter=lfs diff=lfs merge=lfs -text
272
+ components/vulkan/2610100032004/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=lfs -text
273
+ components/vulkan/2610100032004/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
274
+ components/macos/2610100032005/ggml-metal-aarch64.dylib filter=lfs diff=lfs merge=lfs -text
components/cuda/2610100758001/BUILD_INFO.cuda.md ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # cuda 2610100758001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-x86_64.so` | x86_64 | cuda | `6628879f5b518960e37c67f798c8adc756c32274a2d87908c66e0965a9913327` | 756031640 |
10
+ | `ggml-cuda-win-x86_64.dll` | win-x86_64 | cuda | `a3c339d4ed8f4cdec0524faace60efff9df4f0d74dc54edf88443953773093a1` | 730134016 |
11
+
12
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
13
+
14
+ ```
15
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
16
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
17
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
18
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
19
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
20
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
21
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
22
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
23
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
24
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
25
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
26
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
27
+ ```
28
+
29
+ CUDA 13. Six generations, 406 kernel images each; `120` is the sm_120f kernel set. Built at patch chain 0611 from the
30
+ clean-room tree of the 0611 engine (0612 changed no library input that the library reaches) (0611: the library takes the KVarN prompt-read chunk from the host, `ggml_backend_kvarn_set_window_chunk`); NEW `ggml` digest (0590) — the c10 library does not load on this engine. Adds the PQ2_0 /
31
+ PTQ1_0 MMVQ + MMQ kernels, the fused FWHT quantizer, the gated-delta-net kernels (bf16 state), the GB10 / Hopper kernel
32
+ sets, the pair flash-attention loader as the default.
33
+ Gated with the 2610100940001 engine: bs2 (sm_120) Bonsai-2 CUDA + Vulkan gate 56/56 (KLD vs the CUDA base 0.000168 / 0.000311, parity, speed, needle 3/3 at 112k); old-glibc 2.28 image, both backends; fit16 on the RTX 3090 held to 14985 MiB free (the drafter dropped, the window fully resident, 4/4 boots, 54.5–55.5 tok/s; greedy text identical to the 0611 engine on every arm); Windows RTX 3090 with the six-arch DLL 9/9 (MTP 1.20x, vision).
components/cuda/2610100758001/ggml-cuda-win-x86_64.dll ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a3c339d4ed8f4cdec0524faace60efff9df4f0d74dc54edf88443953773093a1
3
+ size 730134016
components/cuda/2610100758001/ggml-cuda-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6628879f5b518960e37c67f798c8adc756c32274a2d87908c66e0965a9913327
3
+ size 756031640
components/cuda12/2610100758002/BUILD_INFO.cuda12.md ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # cuda12 2610100758002
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-cu12-x86_64.so` | x86_64 | cuda12 | `6f186bf7c192a82c35a539886f314e3d2642875c7ea8e4e0cbb20627cf567c2f` | 260609160 |
10
+
11
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
12
+
13
+ ```
14
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
15
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
16
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
17
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
18
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
19
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
20
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
21
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
22
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
23
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
24
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
25
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
26
+ ```
27
+
28
+ CUDA 12.8, for GPUs older than sm_75 (Maxwell, Pascal, Volta). Built at patch chain 0611. **These bytes have not been run on a
29
+ card** (no such GPU reachable at the cut); the forced arm (`OPENCOTI_CUDA_LEGACY=1`, the library beside the binary) is the proof
30
+ of loading. The ternary MMQ kernels need integer dot (sm_61+): on sm_52 the types take the dequant path.
components/cuda12/2610100758002/ggml-cuda-cu12-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6f186bf7c192a82c35a539886f314e3d2642875c7ea8e4e0cbb20627cf567c2f
3
+ size 260609160
components/engine/2610100921001/BUILD_INFO.engine.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # engine 2610100921001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ | file | platform | kind | sha256 | bytes |
6
+ |---|---|---|---|---|
7
+ | `opencoti-0.10.5-c11-2610100921001` | any | bin | `2deabf11d0f6d97618449e2227f7e0b6c179038a82af2fa5cac95ca15acd28a2` | 135607755 |
8
+ | `opencoti-0.10.5-c11-2610100921001.exe` | win-x86_64 | bin | `2deabf11d0f6d97618449e2227f7e0b6c179038a82af2fa5cac95ca15acd28a2` | 135607755 |
9
+
10
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
11
+
12
+ ```
13
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
14
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
15
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
16
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
17
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
18
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
19
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
20
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
21
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
22
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
23
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
24
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
25
+ ```
26
+
27
+ ABI `media` = `b099409b3029a6e380268e31d1caefa6706569961d32e72b3543fbc73903d95c` (exported) — sha256 of this list:
28
+
29
+ ```
30
+ llamafile/oc-audiocpp/audiocpp.h 6ad0b55d5f28adc9ad60a45bf9a7caa428c4f5441389a6932d6670a8e2a90d73
31
+ llamafile/oc-codec/oc_codec_sidecar.h f2c9001d1414aa2a55533b867b79c52a52c682558e0234c5bebd30bf97b91005
32
+ ```
33
+
34
+ ABI `loader` = `d6ff34022f1d3fb5af56ca0c80e3f4e552bf09f1bebdf0cb5f314a454bf38dbc` (exported) — sha256 of this list:
35
+
36
+ ```
37
+ .cosmocc/4.0.2/bin/ape-m1.c 78af79e20abbb3f99a97355959b59bd58949564064adb754f0302ebabb4bf7a3
38
+ ```
39
+
40
+ **c11 dev snapshot — the release candidate of c11 for integration testing.** Engine 2610100724001 (dev build, chain 0001…0611: the 0608 engine every c11 row gate ran on, plus 0609 the c11 tag and 0610 fit-head-exact / 0611 kvarn-chunk-predict — the fix for xollama #939/#940, the drafter on a 16 GB card: KVarN record window 1/500 → 339/500 groups, 3.3 → 47 tok/s at depth on the 3090 held to 14985 MiB free, greedy text unchanged where the window was resident — the c11 release engine is this chain from the clean-room). EVERY GPU library of this index is a NEW full-arch build of the same tree (ggml.h changed in 0590 for the ternary types: a c10 library is refused by name); the Metal library is new; the media sidecars and the macOS loader are the c10 files.
41
+
42
+ What is new since c10 — **Ternary Bonsai 2 27B (PrismML)** is the whole of c11:
43
+
44
+ - **The two ternary weight types** `PQ2_0` (id 142) and `PTQ1_0` (id 143), at PrismML's ids and bytes (0590): a model file
45
+ quantized by the PrismML fork loads unchanged. Every block of the published files dequantizes to the same values as the fork
46
+ and as `gguf-py`; a NaN scale or the legacy group-128 layout under id 42 is refused by name.
47
+ - **The Hadamard weight-fold runtime** (0588): Bonsai 2 stores its weights in a rotated basis (blockwise Hadamard + fixed
48
+ signs); the engine rotates every folded matmul's input at the model's width, natively on CPU, CUDA and Vulkan — never a
49
+ dense rotation matmul. The boot log counts the transforms per graph.
50
+ - **CPU** (0592): `PQ2_0` is repacked at load into the CPU kernels' layout; both types have their dot kernels on x86 and arm64.
51
+ - **CUDA** (0596–0603): MMVQ and MMQ kernels for both types, the fused FWHT + q8 quantizer, the gated-delta-net kernels with
52
+ the bf16 state cache (0603), the DGX-Spark / GB10 (sm_121) and Hopper (sm_90a) kernel sets of the fork; the pair flash-
53
+ attention loader is the default for every in-place KV read (0597 / 0600) and `--fa-quant-prompt auto|in-place|convert`
54
+ picks the quantized-KV prompt route (0601); on sm_86 q4_0 KV prompts read the cache in place (0599).
55
+ - **The qwen35 / gated-delta-net graph** (0602): raw gates, joint q/k normalisation, the split SwiGLU, the rows op —
56
+ launch-for-launch at the fork's count minus the on-device sampler.
57
+ - **Vulkan** (0608): both types on Vulkan — dequant, get_rows, mat-vec, integer-dot mat-vec, mat-mat (scalar + coopmat1;
58
+ no coopmat2 decoder, as the fork), and the Windows `ggml-vulkan.dll`. Vulkan == CUDA greedy except knife-edge margins;
59
+ KLD vs the reference 0.000311 (CUDA 0.000168).
60
+ - **The bundled NextN drafter** (0603–0605): the published MTP bundles carry the drafter head; a ternary target drafts at
61
+ depth 2 by default (the measured best rung: bs2 +10 %, RTX 3090 +16 % over depth 3); the allocator samples a shape only
62
+ from its third consecutive round (bug-3983); the depth ladder is opt-in (`OPENCOTI_MTP_LADDER=1`).
63
+ - **Catalog**: family `bonsai-2` → `ManniX-ITA/Ternary-Bonsai-2-27B-MTP-GGUF` (the MTP bundles, PQ2_0 default and PTQ1_0),
64
+ the optional vision projector (`--mmproj`, loaded lazily at the first image), the per-model KV anchor `kvGBPer32k`
65
+ (2.15 GB / 32K for this GDN hybrid — a 32 GB card is no longer reported "over" at 262K, bug-3987).
66
+ - **A drafter on a card the model nearly fills** (0610 / 0611; xollama #939/#940): the KV window is sized against what the
67
+ assistant head and the SWA ring TAKE — the ring is built first, the server measures the head's take and rebuilds the target's
68
+ memory against it — and the KVarN sizer predicts the prompt-read F16 chunk instead of holding its maximum (the GPU library
69
+ takes the chunk from the host). Gemma-4 A4B + the assistant drafter at 14985 MiB free (an RTX 5080 / 3090 held to it): record
70
+ window 1/500 → 339/500 groups, 3.3 → 47 tok/s at 45k tokens; the sizer WARNs the shortfall when the card is still too full.
71
+ **And when the window still streams with the head attached (0612)**, the server drops the drafter when the head's take covers
72
+ the shortfall — the draft context, the head's weights and its reserve — and rebuilds the target memory once more: on that card
73
+ the record window goes fully resident and the model runs at its no-drafter speed (54–56 tok/s instead of 47 with a streaming
74
+ window); the boot log says `[spec] assistant drafter DROPPED …` or, when dropping would not buy residency, `… KEPT …`. A failed
75
+ compute-buffer re-reserve now fails the decode instead of crashing the process (bug-3995: upstream ignored the reserve's return).
76
+ - **Boot time** (0606 / 0607): the two-pass measure cache no longer pins its host tail (uncapped 128k boot 14.9 → 5.7 s,
77
+ capped 32.9 → 21.6 s); head mode no longer consults the window sizer (bug-3961).
78
+ - **Windows**: a process whose GPU library failed to load no longer reports "Vulkan is not usable" after a crash at exit
79
+ (carried from c9); the home fallback for a bundled library is the user's profile (c10 r3).
80
+ - **Interface**: `ggml.h` changed (the two new types), so the `ggml` digest of this cut is new — every GPU library and the
81
+ Metal library of this index are new builds; a c10 library is refused by the engine with the digest named. The `media` and
82
+ `loader` digests are unchanged: the media sidecars and the macOS loader are the c10 files.
83
+
84
+ Measured (bs2, RTX PRO 6000 Blackwell, sm_120; reference = the PrismML fork b10754 on the same GGUF): PTQ1_0 pp 3353 vs 3441,
85
+ tg 133.4 vs 135.7 off / 160.9 on (depth 2); KLD 0.000168 (the reference's own `-ub 256` floor 0.000163); the FWHT fusion off
86
+ BIT-IDENTICAL to the reference; non-Bonsai models (Omnimerge v6 Q4_K_M, Qwen3-14B Q8_0, Qwen3.6-A3B) bit-identical to c10 on
87
+ CUDA, Vulkan and CPU. Windows 11 (RTX 3090): CUDA MTP on 92.6 vs off 75.9 tok/s (1.22x); Vulkan 11/11, MTP no gain on that card
88
+ (c12 row 30.2). Linux RTX 3090 (220 W cap): MTP 1.10x PQ2_0 / 1.26x PTQ1_0.
89
+
90
+ Known, by design in c11: no coopmat2 (tensor-core) mat-mat for the ternary types on Vulkan (~55–65 % of CUDA on the same
91
+ card; c12 row 30.1); no Metal kernels for the ternary types (they run on the CPU on a Mac, with the WARN naming the type);
92
+ the AMD RX 9070 XT arm is tested by xollama after this cut.
components/engine/2610100921001/opencoti-0.10.5-c11-2610100921001 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2deabf11d0f6d97618449e2227f7e0b6c179038a82af2fa5cac95ca15acd28a2
3
+ size 135607755
components/engine/2610100921001/opencoti-0.10.5-c11-2610100921001.exe ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2deabf11d0f6d97618449e2227f7e0b6c179038a82af2fa5cac95ca15acd28a2
3
+ size 135607755
components/macos/2610100032005/BUILD_INFO.macos.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # macos 2610100032005
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ape-macos-aarch64` | macos-aarch64 | ape | `07f550415a59415973ef877eec99a776dfa5263b06fdb1fd9cefb7b884f0f723` | 56896 |
10
+ | `ggml-metal-aarch64.dylib` | macos-aarch64 | metal | `81b8a7c812d5192067d165d33c02a5cc8fce1a92d0d55b44ecd311127fa4f452` | 1614736 |
11
+
12
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
13
+
14
+ ```
15
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
16
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
17
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
18
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
19
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
20
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
21
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
22
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
23
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
24
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
25
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
26
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
27
+ ```
28
+
29
+ ABI `loader` = `d6ff34022f1d3fb5af56ca0c80e3f4e552bf09f1bebdf0cb5f314a454bf38dbc` (exported) — sha256 of this list:
30
+
31
+ ```
32
+ .cosmocc/4.0.2/bin/ape-m1.c 78af79e20abbb3f99a97355959b59bd58949564064adb754f0302ebabb4bf7a3
33
+ ```
34
+
35
+ Apple notarisation (Accepted), submission id(s): `baf2980f-282d-4bcc-84d3-fb13122d2c7d eea7c66d-e8a4-4860-b9da-71e9416a2805`.
36
+
37
+ The prebuilt APE loader (the c10 file, `loader` digest unchanged) and a NEW prebuilt Metal library, built on the Mac by
38
+ engine 2610100032005 from its embedded sources: it states the new `ggml` digest. No Metal kernels for PQ2_0 / PTQ1_0 (CPU
39
+ fallback with the WARN). Notarised: Metal library baf2980f-282d-4bcc-84d3-fb13122d2c7d, loader eea7c66d-e8a4-4860-b9da-71e9416a2805 (Accepted). Gated on the Mac mini (M-series) in the release form — engine + Metal library + the three media sidecars in one directory, fresh HOME, no compiler on PATH — with the 2610100940001 engine: 21/21 (Metal loaded, 117.6 tok/s, Kokoro + Supertonic speech through the sidecars, the eSpeak-less refusal by name).
components/macos/2610100032005/ape-macos-aarch64 ADDED
Binary file (56.9 kB). View file
 
components/macos/2610100032005/ggml-metal-aarch64.dylib ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:81b8a7c812d5192067d165d33c02a5cc8fce1a92d0d55b44ecd311127fa4f452
3
+ size 1614736
components/media/2610071902001/BUILD_INFO.media.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610071819001; gated floor (engine-min) 2610040950001; gated with 2610071819001, 2610050707001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
@@ -32,6 +32,12 @@ llamafile/oc-codec/oc_codec_sidecar.h f2c9001d1414aa2a55533b867b79c52a52c682558e
32
 
33
  Apple notarisation (Accepted), submission id(s): `83ee464c-3816-4b5c-9da7-cd4313840f25 aa53f9d9-14ba-443a-a059-64f82416f86a`.
34
 
 
 
 
 
 
 
35
  Twelve sidecars, three per platform. **New in this component: the four `oc-audiocpp-*`** (audio.cpp patch 0007, bug-3955); `oc-codec-*` and `oc-espeak-*` are the files of media 2610040945001, unchanged.
36
 
37
  0007: a speech model whose weights are unpacked to disk (Supertonic, Kokoro, KittenTTS) failed for every account but the one that first wrote the shared `/tmp/audiocpp-gguf` (and the eSpeak data the same way). Now the unpacked files go to the user's own cache directory (`$XDG_CACHE_HOME` or `~/.cache`, `~/Library/Caches`, `%LOCALAPPDATA%`, under `audiocpp/`), then to a private per-user temp directory (`<tmp>/audiocpp-gguf-<uid>`, made 0700, refused when it is not the user's own), the old shared directory last; `AUDIOCPP_GGUF_CACHE` overrides. A failure names every directory tried. The `media` interface is unchanged (same `audiocpp.h`), so each file states the same `media` digest and works with the same engines.
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610071819001; gated floor (engine-min) 2610040950001; gated with 2610050707001, 2610071819001, 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
32
 
33
  Apple notarisation (Accepted), submission id(s): `83ee464c-3816-4b5c-9da7-cd4313840f25 aa53f9d9-14ba-443a-a059-64f82416f86a`.
34
 
35
+ The twelve sidecars of c10, unchanged (same `media` digest): oc-codec, oc-audiocpp (0007), oc-espeak, three per platform.
36
+
37
+ ---
38
+
39
+ The c10 notes:
40
+
41
  Twelve sidecars, three per platform. **New in this component: the four `oc-audiocpp-*`** (audio.cpp patch 0007, bug-3955); `oc-codec-*` and `oc-espeak-*` are the files of media 2610040945001, unchanged.
42
 
43
  0007: a speech model whose weights are unpacked to disk (Supertonic, Kokoro, KittenTTS) failed for every account but the one that first wrote the shared `/tmp/audiocpp-gguf` (and the eSpeak data the same way). Now the unpacked files go to the user's own cache directory (`$XDG_CACHE_HOME` or `~/.cache`, `~/Library/Caches`, `%LOCALAPPDATA%`, under `audiocpp/`), then to a private per-user temp directory (`<tmp>/audiocpp-gguf-<uid>`, made 0700, refused when it is not the user's own), the old shared directory last; `AUDIOCPP_GGUF_CACHE` overrides. A failure names every directory tried. The `media` interface is unchanged (same `audiocpp.h`), so each file states the same `media` digest and works with the same engines.
components/sbsa/2610100758003/BUILD_INFO.sbsa.md ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # sbsa 2610100758003
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-cuda-sbsa-aarch64.so` | aarch64 | cuda | `ff733a2f9e6aa6d27ca3febb05b632fe1e5ef467668763afb108a56396488ef6` | 346864920 |
10
+
11
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
12
+
13
+ ```
14
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
15
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
16
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
17
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
18
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
19
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
20
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
21
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
22
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
23
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
24
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
25
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
26
+ ```
27
+
28
+ CUDA 13 for Linux aarch64 (SBSA: DGX Spark / GB10 = sm_121, Jetson Thor). Built at patch chain 0611; carries the fork's GB10
29
+ mmvq table and exp(g) pre-pass (sm_121). **Never run: there is no such machine in the test fleet.**
components/sbsa/2610100758003/ggml-cuda-sbsa-aarch64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ff733a2f9e6aa6d27ca3febb05b632fe1e5ef467668763afb108a56396488ef6
3
+ size 346864920
components/vulkan/2610100032004/BUILD_INFO.vulkan.md ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # vulkan 2610100032004
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
+
7
+ | file | platform | kind | sha256 | bytes |
8
+ |---|---|---|---|---|
9
+ | `ggml-vulkan-x86_64.so` | x86_64 | vulkan | `449b4cca40b0c056b369b6887c486e2a6509ba8d88c8fe2adabb5bbc06b9c070` | 61269328 |
10
+ | `ggml-vulkan-win-x86_64.dll` | win-x86_64 | vulkan | `10076c3f3b4c304c52bffe3bd02c603a5fd5001f33feea3786c23cff839d868a` | 62313243 |
11
+
12
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) — sha256 of this list:
13
+
14
+ ```
15
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
16
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
17
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
18
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
19
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
20
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
21
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
22
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
23
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
24
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
25
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
26
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
27
+ ```
28
+
29
+ Built at patch chain 0609 (0608: the ternary types on Vulkan); the same bytes serve the 0612 engine — no Vulkan input changed in
30
+ 0610 / 0611 / 0612 and the `ggml` digest is unchanged. Gated with the 2610100940001 engine: bs2 Vulkan cases of the 56/56 gate (PQ2_0 / PTQ1_0 KLD 0.000311 vs the CUDA base, needle 3/3 at 112k, tg 81.8 / 83.5 t/s); old-glibc 2.28 image; Windows RTX 3090 Vulkan 11/11 (NextN 560 draft steps on Vulkan). AMD Radeon on Windows: not run by us at this
31
+ cut (xollama, RX 9070 XT, after the cut).
components/vulkan/2610100032004/ggml-vulkan-win-x86_64.dll ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10076c3f3b4c304c52bffe3bd02c603a5fd5001f33feea3786c23cff839d868a
3
+ size 62313243
components/vulkan/2610100032004/ggml-vulkan-x86_64.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:449b4cca40b0c056b369b6887c486e2a6509ba8d88c8fe2adabb5bbc06b9c070
3
+ size 61269328