ManniX-ITA commited on
Commit
ce78e34
Β·
verified Β·
1 Parent(s): 262af96

pin v2 payload: engine 2610101319001, cuda 2610100758001, cuda12 2610100758002, sbsa 2610100758003, vulkan 2610100032004, macos 2610100032005, media 2610071902001

Browse files
.gitattributes CHANGED
@@ -272,3 +272,5 @@ components/sbsa/2610100758003/ggml-cuda-sbsa-aarch64.so filter=lfs diff=lfs merg
272
  components/vulkan/2610100032004/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=lfs -text
273
  components/vulkan/2610100032004/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
274
  components/macos/2610100032005/ggml-metal-aarch64.dylib filter=lfs diff=lfs merge=lfs -text
 
 
 
272
  components/vulkan/2610100032004/ggml-vulkan-x86_64.so filter=lfs diff=lfs merge=lfs -text
273
  components/vulkan/2610100032004/ggml-vulkan-win-x86_64.dll filter=lfs diff=lfs merge=lfs -text
274
  components/macos/2610100032005/ggml-metal-aarch64.dylib filter=lfs diff=lfs merge=lfs -text
275
+ components/engine/2610101319001/opencoti-0.10.5-c11-2610101319001 filter=lfs diff=lfs merge=lfs -text
276
+ components/engine/2610101319001/opencoti-0.10.5-c11-2610101319001.exe filter=lfs diff=lfs merge=lfs -text
components/cuda/2610100758001/BUILD_INFO.cuda.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
components/cuda12/2610100758002/BUILD_INFO.cuda12.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
components/engine/2610101319001/BUILD_INFO.engine.md ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # engine 2610101319001
2
+
3
+ Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
+
5
+ | file | platform | kind | sha256 | bytes |
6
+ |---|---|---|---|---|
7
+ | `opencoti-0.10.5-c11-2610101319001` | any | bin | `2a2f9eeffaabaab48c05f1f077b9318d9fe21185bef2afb9758cccd8626d91ff` | 135607617 |
8
+ | `opencoti-0.10.5-c11-2610101319001.exe` | win-x86_64 | bin | `2a2f9eeffaabaab48c05f1f077b9318d9fe21185bef2afb9758cccd8626d91ff` | 135607617 |
9
+
10
+ ABI `ggml` = `ea73e6b2e6fba55c93a75ee8f3e5ba968e906eb20faf8746c792ead5c92f1285` (exported) β€” sha256 of this list:
11
+
12
+ ```
13
+ llama.cpp/ggml/include/ggml-alloc.h 94e4cd069b9313b2ceb35dacec901981e0bb478d8bb31035b7126be091998c23
14
+ llama.cpp/ggml/include/ggml-backend.h 764d1640ad62579513c4837539708f89e8603c9dbeae053a13d1cbd84a82c865
15
+ llama.cpp/ggml/include/ggml-cuda.h 674ea064021cf5b9970e78f1209e513dd23350b94a44a705a27eec1423fe4b0a
16
+ llama.cpp/ggml/include/ggml-metal.h 322f36cd30f3e9e7aad7b5b9bc63078012fd0b7706ac4e24f5721b604f3d8980
17
+ llama.cpp/ggml/include/ggml-neo-pipeline.h 190f39e183542d16588a3104d8cb4359bf52fc7a62afe8d8cc4c822f4085086d
18
+ llama.cpp/ggml/include/ggml-oc-kvarn-range.h 76635fdd8978632f670640ddf5125a666441e3d47c680bb8b5c0a099e4861202
19
+ llama.cpp/ggml/include/ggml-oc-ledger.h af6dc423ee7c6c3f67a6a187d6dfbfeb94e52ddc1db27c45ee7440289bc44282
20
+ llama.cpp/ggml/include/ggml-vulkan.h 7eae5dad2cc7bb4d3eca828539f441816d5fd59fdc7b224d49aa229fc7b7248c
21
+ llama.cpp/ggml/include/ggml.h 2fcd3834a184797b0f105429c36b246a4100ad30693754d44becd1a36c17ea68
22
+ llama.cpp/ggml/include/opencoti-devtype.h 94dea946173449bb76bffd7f733ab822fe658fe7198992de750ac012d29308ec
23
+ llama.cpp/ggml/src/ggml-backend-impl.h f9fdfe6bae61d1387593b40fc0f90ceee490c6a339ec6d7388c51e4fca318d8f
24
+ llamafile/gpu_backend.h d17c504fbab50ff9376105a1f2c79d99fcde4b431886b17b90718a2a359278de
25
+ ```
26
+
27
+ ABI `media` = `b099409b3029a6e380268e31d1caefa6706569961d32e72b3543fbc73903d95c` (exported) β€” sha256 of this list:
28
+
29
+ ```
30
+ llamafile/oc-audiocpp/audiocpp.h 6ad0b55d5f28adc9ad60a45bf9a7caa428c4f5441389a6932d6670a8e2a90d73
31
+ llamafile/oc-codec/oc_codec_sidecar.h f2c9001d1414aa2a55533b867b79c52a52c682558e0234c5bebd30bf97b91005
32
+ ```
33
+
34
+ ABI `loader` = `d6ff34022f1d3fb5af56ca0c80e3f4e552bf09f1bebdf0cb5f314a454bf38dbc` (exported) β€” sha256 of this list:
35
+
36
+ ```
37
+ .cosmocc/4.0.2/bin/ape-m1.c 78af79e20abbb3f99a97355959b59bd58949564064adb754f0302ebabb4bf7a3
38
+ ```
39
+
40
+ **c11 dev snapshot r2 β€” engine 2610101319001 (dev build, chain through 0613-gdn-nonfused-cont, the bug-3999 fix): the previous snapshot's engine 2610100921001 aborted at load on Qwen3.5 0.8B–4B whenever the memory fit streamed a host tail; the libraries are unchanged.** (The previous snapshot's note named engine 2610100724001 by mistake; its bytes were 2610100921001.)
41
+
42
+ **c11 dev snapshot β€” the release candidate of c11 for integration testing.** Engine 2610100724001 (dev build, chain 0001…0611: the 0608 engine every c11 row gate ran on, plus 0609 the c11 tag and 0610 fit-head-exact / 0611 kvarn-chunk-predict β€” the fix for xollama #939/#940, the drafter on a 16 GB card: KVarN record window 1/500 β†’ 339/500 groups, 3.3 β†’ 47 tok/s at depth on the 3090 held to 14985 MiB free, greedy text unchanged where the window was resident β€” the c11 release engine is this chain from the clean-room). EVERY GPU library of this index is a NEW full-arch build of the same tree (ggml.h changed in 0590 for the ternary types: a c10 library is refused by name); the Metal library is new; the media sidecars and the macOS loader are the c10 files.
43
+
44
+ What is new since c10 β€” **Ternary Bonsai 2 27B (PrismML)** is the whole of c11:
45
+
46
+ - **The two ternary weight types** `PQ2_0` (id 142) and `PTQ1_0` (id 143), at PrismML's ids and bytes (0590): a model file
47
+ quantized by the PrismML fork loads unchanged. Every block of the published files dequantizes to the same values as the fork
48
+ and as `gguf-py`; a NaN scale or the legacy group-128 layout under id 42 is refused by name.
49
+ - **The Hadamard weight-fold runtime** (0588): Bonsai 2 stores its weights in a rotated basis (blockwise Hadamard + fixed
50
+ signs); the engine rotates every folded matmul's input at the model's width, natively on CPU, CUDA and Vulkan β€” never a
51
+ dense rotation matmul. The boot log counts the transforms per graph.
52
+ - **CPU** (0592): `PQ2_0` is repacked at load into the CPU kernels' layout; both types have their dot kernels on x86 and arm64.
53
+ - **CUDA** (0596–0603): MMVQ and MMQ kernels for both types, the fused FWHT + q8 quantizer, the gated-delta-net kernels with
54
+ the bf16 state cache (0603), the DGX-Spark / GB10 (sm_121) and Hopper (sm_90a) kernel sets of the fork; the pair flash-
55
+ attention loader is the default for every in-place KV read (0597 / 0600) and `--fa-quant-prompt auto|in-place|convert`
56
+ picks the quantized-KV prompt route (0601); on sm_86 q4_0 KV prompts read the cache in place (0599).
57
+ - **The qwen35 / gated-delta-net graph** (0602): raw gates, joint q/k normalisation, the split SwiGLU, the rows op β€”
58
+ launch-for-launch at the fork's count minus the on-device sampler.
59
+ - **Vulkan** (0608): both types on Vulkan β€” dequant, get_rows, mat-vec, integer-dot mat-vec, mat-mat (scalar + coopmat1;
60
+ no coopmat2 decoder, as the fork), and the Windows `ggml-vulkan.dll`. Vulkan == CUDA greedy except knife-edge margins;
61
+ KLD vs the reference 0.000311 (CUDA 0.000168).
62
+ - **The bundled NextN drafter** (0603–0605): the published MTP bundles carry the drafter head; a ternary target drafts at
63
+ depth 2 by default (the measured best rung: bs2 +10 %, RTX 3090 +16 % over depth 3); the allocator samples a shape only
64
+ from its third consecutive round (bug-3983); the depth ladder is opt-in (`OPENCOTI_MTP_LADDER=1`).
65
+ - **Catalog**: family `bonsai-2` β†’ `ManniX-ITA/Ternary-Bonsai-2-27B-MTP-GGUF` (the MTP bundles, PQ2_0 default and PTQ1_0),
66
+ the optional vision projector (`--mmproj`, loaded lazily at the first image), the per-model KV anchor `kvGBPer32k`
67
+ (2.15 GB / 32K for this GDN hybrid β€” a 32 GB card is no longer reported "over" at 262K, bug-3987).
68
+ - **A drafter on a card the model nearly fills** (0610 / 0611; xollama #939/#940): the KV window is sized against what the
69
+ assistant head and the SWA ring TAKE β€” the ring is built first, the server measures the head's take and rebuilds the target's
70
+ memory against it β€” and the KVarN sizer predicts the prompt-read F16 chunk instead of holding its maximum (the GPU library
71
+ takes the chunk from the host). Gemma-4 A4B + the assistant drafter at 14985 MiB free (an RTX 5080 / 3090 held to it): record
72
+ window 1/500 β†’ 339/500 groups, 3.3 β†’ 47 tok/s at 45k tokens; the sizer WARNs the shortfall when the card is still too full.
73
+ **And when the window still streams with the head attached (0612)**, the server drops the drafter when the head's take covers
74
+ the shortfall β€” the draft context, the head's weights and its reserve β€” and rebuilds the target memory once more: on that card
75
+ the record window goes fully resident and the model runs at its no-drafter speed (54–56 tok/s instead of 47 with a streaming
76
+ window); the boot log says `[spec] assistant drafter DROPPED …` or, when dropping would not buy residency, `… KEPT …`. A failed
77
+ compute-buffer re-reserve now fails the decode instead of crashing the process (bug-3995: upstream ignored the reserve's return).
78
+ - **Boot time** (0606 / 0607): the two-pass measure cache no longer pins its host tail (uncapped 128k boot 14.9 β†’ 5.7 s,
79
+ capped 32.9 β†’ 21.6 s); head mode no longer consults the window sizer (bug-3961).
80
+ - **Windows**: a process whose GPU library failed to load no longer reports "Vulkan is not usable" after a crash at exit
81
+ (carried from c9); the home fallback for a bundled library is the user's profile (c10 r3).
82
+ - **Interface**: `ggml.h` changed (the two new types), so the `ggml` digest of this cut is new β€” every GPU library and the
83
+ Metal library of this index are new builds; a c10 library is refused by the engine with the digest named. The `media` and
84
+ `loader` digests are unchanged: the media sidecars and the macOS loader are the c10 files.
85
+
86
+ Measured (bs2, RTX PRO 6000 Blackwell, sm_120; reference = the PrismML fork b10754 on the same GGUF): PTQ1_0 pp 3353 vs 3441,
87
+ tg 133.4 vs 135.7 off / 160.9 on (depth 2); KLD 0.000168 (the reference's own `-ub 256` floor 0.000163); the FWHT fusion off
88
+ BIT-IDENTICAL to the reference; non-Bonsai models (Omnimerge v6 Q4_K_M, Qwen3-14B Q8_0, Qwen3.6-A3B) bit-identical to c10 on
89
+ CUDA, Vulkan and CPU. Windows 11 (RTX 3090): CUDA MTP on 92.6 vs off 75.9 tok/s (1.22x); Vulkan 11/11, MTP no gain on that card
90
+ (c12 row 30.2). Linux RTX 3090 (220 W cap): MTP 1.10x PQ2_0 / 1.26x PTQ1_0.
91
+
92
+ Known, by design in c11: no coopmat2 (tensor-core) mat-mat for the ternary types on Vulkan (~55–65 % of CUDA on the same
93
+ card; c12 row 30.1); no Metal kernels for the ternary types (they run on the CPU on a Mac, with the WARN naming the type);
94
+ the AMD RX 9070 XT arm is tested by xollama after this cut.
components/engine/2610101319001/opencoti-0.10.5-c11-2610101319001 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2a2f9eeffaabaab48c05f1f077b9318d9fe21185bef2afb9758cccd8626d91ff
3
+ size 135607617
components/engine/2610101319001/opencoti-0.10.5-c11-2610101319001.exe ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2a2f9eeffaabaab48c05f1f077b9318d9fe21185bef2afb9758cccd8626d91ff
3
+ size 135607617
components/macos/2610100032005/BUILD_INFO.macos.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
components/media/2610071902001/BUILD_INFO.media.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610071819001; gated floor (engine-min) 2610040950001; gated with 2610050707001, 2610071819001, 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610071819001; gated floor (engine-min) 2610040950001; gated with 2610050707001, 2610071819001, 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
components/sbsa/2610100758003/BUILD_INFO.sbsa.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
components/vulkan/2610100032004/BUILD_INFO.vulkan.md CHANGED
@@ -2,7 +2,7 @@
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
- Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|
 
2
 
3
  Component of the opencoti pin format v2 (docs/protocols/PIN_FORMAT.md).
4
 
5
+ Built from the engine source of build 2610100940001; gated floor (engine-min) 2610100921001; gated with 2610100921001, 2610101319001.
6
 
7
  | file | platform | kind | sha256 | bytes |
8
  |---|---|---|---|---|