Download docs/PERFORMANCE.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 26.1 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/PERFORMANCE.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/PERFORMANCE.md
-
curl -L -o PERFORMANCE.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/PERFORMANCE.md
bankml β performance
Every number in this file was measured, on the hardware named beside it, and can be reproduced with the command
given. Nothing is quoted from upstream. Where a result is a lower bound, an estimate or a single run, it says so.
The per-release records are in CHANGELOG.md and testing/results/<version>.txt; the kernels'
own pages are q1_0.rs, q2_0.rs, f16.rs, par.rs
and forward.rs.
- Reference engine: llama.cpp b11192, the ubuntu-x64 release checked against its sha256
(
34cf6fa5de9da0db3932c78fe15fed2fbca17451e665dac0a4f6a3c8fc881ec7). Kernel comparisons call that release's own exported symbols in-process, on the same bytes, so the A/B measures code and nothing else. - Models: PrismML Bonsai, forked with provenance on Hugging Face β
Bonsai-8B Q1_0 (1-bit, 1.16 GB, sha256
284a335aβ¦bd54) and Ternary-Bonsai-8B Q2_0_g64 (1.58-bit, 2.31 GB, sha256e17b298dβ¦c55e); the Q1_0 oracle uses Bonsai-1.7B-Q1_0 (sha2563d7c6c90β¦).
Machines
| name | CPU | threads used | RAM | notes |
|---|---|---|---|---|
| laptop | AMD Ryzen 3 3200U (Zen+) | 1 (kernels) Β· 3 (llama-server) | 5.8 GB | the development box; AVX2, no AVX-512 |
| VPS | AMD EPYC 7543P (Zen3), 2 vCPU | 1 (-t 1, CPUQuota 100 %) |
7.8 GB | mindx.pythai.net |
| Space | Hugging Face zero-a10g Space, CPU path |
~16 busy (CPU s Γ· wall s) | container | llama.cpp on CPU; the GPU is unused |
Q2_0_g64 (ternary) kernel against llama.cpp b11192 β laptop, 1 thread
Why ternary is slow in llama.cpp b11192: it has no x86 kernel for Q2_0. On x86 the generic C function is
renamed to ggml_vec_dot_q2_0_q8_0 (ggml/src/ggml-cpu/arch-fallback.h), and there is no repack or sgemm case for the
type; the disassembly of the shipped libggml-cpu-haswell.so is scalar code with 64 imul per block, used for decode
and prefill alike. It costs about 49 ns per 64 weights β roughly 7Γ what ggml spends per weight on Q1_0.
BANKML_GGML_LIB=<b11192 release dir> cargo test --release -- --ignored ab_vs_ggml_q2_0 --nocapture --test-threads=1
β real layer-0 weights of Ternary-Bonsai-8B, the two kernels alternated in one process on the same bits.
| matmul (rows Γ columns) | llama.cpp b11192, ns/block min / median | bankml, ns/block min / median | speed-up (median) |
|---|---|---|---|
| attn_q 4096 Γ 4096 | 48.98 / 49.44 | 4.98 / 5.15 | 9.60Γ |
| attn_k 1024 Γ 4096 | 48.63 / 48.97 | 4.89 / 4.98 | 9.83Γ |
| attn_output 4096 Γ 4096 | 48.93 / 49.92 | 4.98 / 5.13 | 9.73Γ |
| ffn_gate / ffn_up 12288 Γ 4096 | 49.08 / 49.39 | 5.04 / 5.19 | 9.52Γ |
| ffn_down 4096 Γ 12288 | 48.80 / 49.46 | 5.01 / 5.17 | 9.57Γ |
| output 151669 Γ 4096 | 49.61 / 50.73 (492 ms) | 5.11 / 5.29 (51 ms) | 9.59Γ |
| prefill 1024 rows Γ 32 columns, per rowΒ·column | 48.90 / 49.77 (per pair) | 3.96 / 3.97 (1Γ4 tile) | 12.54Γ |
Over four runs: decode 9.1β10.6Γ, prefill tile 12.3β13.1Γ. Absolute ns move with the laptop's boost clock (ggml read 34β55 ns across runs); both kernels move together, so the ratio is the result. Preparing the activation once per token costs 11 Β΅s (n = 4096) and 29 Β΅s (n = 12288), about 2 ms per token β 0.5 % of the matmul time, not included above. From L1 the kernel runs at about 4.0 ns/block against a measured memory floor of about 2.3 ns (7.7 GB/s single-thread read), so it is still partly compute-bound.
Exactness (oracle_ggml_b11192_real_ternary_bonsai_8b, re-run independently on 2026-09-26): all 254 Q2_0
tensors β 8,188,239,872 weights β dequantized bit-exact; 762 / 762 q8_0 activation rows byte-exact; 762 / 762
dot products bit-exact against ggml's haswell build (scalar model, AVX2 and portable paths), and a no-FMA model equal to
ggml's baseline libggml-cpu-x64.so in 762 / 762. The haswell and x64 builds themselves disagree on 19 of the 762 β the
oracle does test float order. The format's codes are {β1, 0, +1, +2}; in the real file +2 never occurs (31 % / 38 % /
31 % / 0 %).
Q1_0 (1-bit) kernel against llama.cpp b11192 β laptop, 1 thread
BANKML_GGML_LIB=<b11192 release dir> cargo test --release -- --ignored --nocapture --test-threads=1
| measurement | bankml | llama.cpp b11192 | ratio |
|---|---|---|---|
| decode GEMV 12288Γ4096, ns/block, min | 14.32 (mat_vec) |
14.43 | 1.01Γ |
| decode GEMV, ns/block, median | 14.47 | 14.93 | 1.03Γ |
decode GEMV via vec_dot_act (activation scales converted once per token), min / median |
14.37 / 14.61 | 14.43 / 14.93 | 1.00Γ / 1.02Γ |
decode GEMV via per-pair vec_dot, min / median |
15.64 / 16.00 | 14.43 / 14.93 | 0.92Γ / 0.93Γ |
| prefill 1024Γ4096 Γ 32 columns, ns/(blockΒ·column), min / median | 12.95 / 13.31 (1Γ4 tile) | 14.22 / 15.43 (per-pair) | 1.10Γ / 1.16Γ |
Within the crate: the AVX2 path is 15.2Γ the portable C port of ggml's generic kernel and 20.1Γ the scalar reference model (L1-resident row, n = 4096: 16.25 vs 246.28 vs 327.02 ns/block).
Exactness (oracle_ggml_b11192_real_bonsai_1_7b, against the release's libggml-cpu-haswell.so):
all 197 Q1_0 tensors β 1,719,904,256 weights β dequantized bit-exact; 788 / 788 q8_0 activation rows
byte-exact; 788 / 788 dot products bit-exact against ggml's AVX2 kernel and against its generic kernel.
End-to-end baselines: llama.cpp b11192 serving Bonsai (the bar bankml must meet)
Tokens are llama-server's own counts. "Read" is prompt evaluation; "write" is generation.
1-bit Bonsai-8B
| where | read tok/s | write tok/s | representative run |
|---|---|---|---|
| laptop, 3 threads | 2.3β3.2 | 1.5β2.4 (0.43 under memory pressure) | 310-token prompt: 137 s cold, ~7 s cached |
| VPS, 1 thread | 2.1β3.1 | 1.1β2.5 | 5 hourly runs, 272β299 tokens in 222β585 s |
| Space | cached 310-token prompt in 0.3β0.6 s | 23β26 (warm) | 198 tokens in 7.9β8.5 s |
Ternary-Bonsai-8B vs 1-bit β the same question under the same system prompt (314 prompt tokens)
| where | model | state | first token | total | written | rate |
|---|---|---|---|---|---|---|
| Space | 1-bit | warm | 0.05 s | 5.0 s | 138 | ~28 tok/s |
| Space | Ternary | warm | 0.3 s | 21.8 s | 82 | ~3.8 tok/s |
| laptop | Ternary | cold | 638.7 s | 868.6 s | 94 | read 0.49, write 0.40 tok/s |
| laptop | Ternary | warm | 2.6 s | 214.6 s | 89 | write 0.42 tok/s |
Ternary runs about 5β7Γ slower than 1-bit on the same CPU under llama.cpp b11192 (laptop ~0.4 vs ~2 tok/s; Space ~3.8 vs ~26 tok/s), although its file is only 2Γ larger. The next two sections show why, and what bankml's ternary kernel does about it.
Whole-model decode budget β where one ternary token's time goes
decode_budget_q2_0: one token's worth of all 253 ternary matmuls (2.13 GB of weights), each tensor once in file
order as a real token touches them; llama-bench run immediately before and after each pass, clock logged.
| laptop | llama.cpp end-to-end (llama-bench) | llama.cpp matmuls only | bankml matmuls only | bankml ceiling (matmuls only) |
|---|---|---|---|---|
| 3 threads (3.2β3.3 GHz) | 0.386 / 0.373 tok/s = 2.59 / 2.68 s/token | 2.54 / 2.64 s | 0.36 / 0.37 s (7.0Γ) | 2.77 tok/s |
| 1 thread (2.78β3.10 GHz) | 4.50 / 5.21 s/token | 4.22 / 4.42 s | 0.49 s (8.7Γ) | 2.06 tok/s |
- The matmul is the wall: about 98 % of llama.cpp's time per token at 3 threads (85β98 % at 1 thread). llama-bench's 3-thread rate reproduces the 0.42 tok/s measured serving the model.
- What it implies: one thread of bankml's ternary matmuls (0.49 s) is about the time llama.cpp needs for a whole 1-bit token on three threads (~2 tok/s) β before attention, normalisation and the rest of the forward pass, which bankml does not have yet (phase P3), so the ceiling is a matmul bound, not an end-to-end rate.
- bankml gains little from threads here (1.34Γ from 1 to 3) β the signature of a memory-bandwidth-bound kernel on a 2-core / 4-thread part. (0.0.2 harness, before the thread pool; superseded by the 0.0.3 section below, which finds both kernels compute-bound.)
- Method notes: unprivileged
perfis blocked on this machine (paranoid = 4), so the matmul share comes from same-clock bracketing, not a cycle profile. A second 1-thread sequence was disturbed by other load and is not used.
Threads (0.0.3) β one token's matmuls at 1β4 threads, both formats
decode_budget_q1_0 / decode_budget_q2_0: all 253 matmuls of one token, each tensor once in file order. ggml's own
kernel and bankml's run on the same pool and row scheduler (par::Pool, 16-row chunks), so the A/B compares kernel
with kernel. Laptop, two runs, min s/token; the full output is in testing/results/0.0.3.txt.
| threads | ternary ggml | ternary bankml | 1-bit ggml | 1-bit bankml |
|---|---|---|---|---|
| 1 | 4.22β4.65 | 0.45β0.47 | 0.62β0.67 | 0.61β0.62 |
| 2 | 2.65 | 0.31 | 0.45β0.46 | 0.43β0.44 |
| 3 | 2.27β2.36 | 0.23β0.25 | 0.34β0.35 | 0.34β0.35 |
| 4 | 2.11 | 0.23 | 0.34β0.35 | 0.36 |
These whole-token figures were measured with the model resident in memory, and reproduced on 2026-09-29 on current code: 0.231 s at three threads, 9.45Γ the reference (below). When other applications leave too little memory to keep the 2.31 GB file cached, the gates measure the disk instead (4.5β5.2 s).
- Ternary is now cheaper than 1-bit. At 3 threads bankml's ternary matmuls (0.23β0.25 s) take less time than ggml's 1-bit ones (0.34β0.35 s), although the ternary weights are twice the bytes. The ternary matmul-only ceiling is 4.0β4.4 tok/s, against 0.42 for llama.cpp.
- The persistent pool replaced per-call thread spawning (12.6 against 88.8 Β΅s per matmul;
bench_pool_overhead). With it, the 3-thread ternary budget moved from 0.36 s (0.0.2 harness) to 0.23β0.25 s. - The 1-bit kernel is at parity at every thread count (0.93β1.13Γ).
- Exactness: the new
oracle_ggml_b11192_real_bonsai_8b_q1_0covers all 254Q1_0tensors of Bonsai-8B (8,188,239,872 weights) bit-exact, 762/762 rows and dot products. Every threaded output in the budgets is compared bit for bit with ggml.
The memory floor (0.0.4) β how far each kernel is from streaming speed
bench_memory_floor: a 768 MiB buffer read on the same pool and scheduler as the matmuls. Laptop:
| threads | read GB/s | floor, ternary token (2.13 GB) | floor, 1-bit token (1.06 GB) |
|---|---|---|---|
| 1 | 14.9β15.2 | 0.14 s | 0.07 s |
| 2β4 | 16.6β17.2 | 0.12β0.13 s | 0.06 s |
bankml's ternary matmuls (0.23β0.27 s at 3 threads) are about 2Γ above the floor; its 1-bit ones (0.34β0.37 s), like ggml's, about 5.5Γ. Both are compute-bound. (0.0.1 quoted 7.7 GB/s single-thread from a scalar loop; this loop vectorises, which is the fair floor.)
1-bit prefill over prepared activations (0.0.4)
ab_vs_ggml (1024 rows Γ 4096 Γ 32 columns, ns per blockΒ·column, min / median):
| kernel | ns | vs ggml per-pair (median) |
|---|---|---|
ggml b11192 per-pair vec_dot |
10.09 / 11.12 | 1.00Γ |
| bankml 1Γ4 tile over q8 bytes (0.0.1) | 9.18 / 9.42 | 1.18Γ |
bankml mat_mul_act, 1Γ4 selection tile over Q8Act |
8.25 / 8.35 | 1.33Γ |
The gate run, disturbed by other load, measured 1.27Γ. Two decode experiments were bit-exact and not faster (see CHANGELOG 0.0.4); 1-bit decode remains at parity with ggml.
Ternary kernel experiments (0.0.5) β laptop, real Ternary-Bonsai-8B tensors
| experiment | bits | speed vs the file-layout kernel | kept |
|---|---|---|---|
Q2Packed: quads repacked (codes contiguous, 4 scales adjacent), same 72 bytes |
exact | 1.00β1.075Γ first run, 0.90β1.05Γ in the gate: noise | no |
| two-row decode tile (activation shared) | exact | 0.80Γ | no |
| software prefetch 256 / 512 / 1024 bytes ahead | exact | Β±3 % | no |
At 3 threads the kernel moves about 9 GB/s against the 17 GB/s floor, and scales 1.8Γ from 1 to 3 threads on two
physical cores. It is bound by compute per core, and at its instruction-throughput limit on Zen+. No kernel changed
in 0.0.5; the three experiments' code is in testing/experiments/.
Reproduce
cargo test --release # unit tests
BANKML_GGML_LIB=/path/to/llama-b11192 testing/release_gate.sh # everything, recorded in testing/results/
BANKML_GGML_LIB=/path/to/llama-b11192 \
cargo test --release -- --ignored --nocapture --test-threads=1 # oracles + A/B + benchmarks
python3 testing/ggml_oracle.py MODEL.gguf /path/to/llama-b11192 .models/oracle-ternary # write an oracle
python3 testing/guard_agree.py target/release/bankml [FILE.gguf ...] # Rust guard == Python guard
The oracle and A/B tests need the model files under .models/ (fetched from the PYTHAI forks above and checked
against their sha256) and the b11192 release directory. They are #[ignore]d so a plain cargo test stays offline.
0.1.8 β 0.2.0: the gateway, the pin, restarts, and what the laptop can hold
Measured on the same Ryzen 3 3200U, on fixed resources where stated (testing/pinned.sh: cores pinned, memory capped
in a cgroup, load recorded).
| what | before | after | how measured |
|---|---|---|---|
| sha256 pin of a 248 MB model (warm cache) | 1.25 s (portable) | 0.23 s (SHA-NI), 5.5Γ | interleaved, 3 runs each (0.1.8) |
| sha256 pin of the 1.16 GB 8B model | β | 2.9 s, equal to sha256sum and the published pin |
0.1.8 |
| carrier restart (hash + load), 8B | 25 s (0.1.7) | 7.0 s | _start_carrier, 0.1.8 |
| first answer after an engine restart (Savante's 282-token system prompt, same question, temperature 0) | 132.3 s (prefill 118 s, 315 tokens) | 15.4 s (restore 0.05 s, prefill 1.2 s, 1 token), identical answer | slot save / restore, 0.1.9 |
| the gateway's own cost per request | β | 0.71 ms (/health 1.08 ms through serve vs 0.37 ms direct) |
200 requests, 0.1.9 |
| cold vs warm hashing (1.7B, pinned) | cold 0.68β0.72 s (364 MB/s) | warm 0.59β0.78 s (419 MB/s) | posix_fadvise(DONTNEED) before each cold run; no read-ahead hint adopted |
| n-gram speculation (8B, 3 prompts Γ 64 tokens, pinned 2 cores, 2.5 GB) | baseline 0.36β1.40 tok/s | ahead in 5 of 6 pairs, +0.01 to +0.36; one β0.06 | token-identical; not beyond this laptop's noise β opt-in |
| draft-model speculation (Bonsai-1.7B for 8B) | 0.47 tok/s | 0.39 / 0.53 | token-identical; no gain β not adopted |
| bge-m3 embedding (Ollama) | β | first call 9.2 s (load), then 1.40 s for 3 texts; 1.14 GB while loaded | 0.1.9 |
the Q2_0 drop-in for llama.cpp (upstream/) |
ggml shipped scalar 0.461 s | 0.134 s (3.4Γ) per 200,000 Γ 4096-wide rows | bit-exact 200,000/200,000, pinned (0.2.0) |
The whole-token ternary budget on this laptop today. The gates since 0.1.0 record 4.5β5.2 s per ternary token at three threads for bankml, and 1.1β1.5Γ the reference, not the 0.23β0.25 s of the 0.0.3β0.0.6 records. The reason was measured on 2026-09-29. With the chat engine resident and other applications holding memory, the 2.31 GB ternary file does not stay in the page cache: 0.49 GB was resident, and after a full read 1.04 GB. Every token then re-reads most of the weights from disk, and both runtimes wait on it. The per-matmul A/B on cached tensors still shows 9.8Γ (0.1.8 gate). The whole-token figure needs a machine with at least 3 GB free to be re-measured; that was done on 2026-09-29 (below).
0.2.1 gate, with the chat engine stopped (2.0 GB free): bankml 1.42 s per ternary token at one and three threads, against llama.cpp's 5.15 s (one thread) and 2.99 s (three threads): 3.6Γ and 2.1Γ. The more of the model stays cached, the closer the whole-token figure comes to the kernels' own ratio.
Re-measured with the model resident (2026-09-29, after 0.2.1)
The browser and the chat engine were closed to free memory, and the ternary file was read twice. 2.25 of its 2.31 GB
stayed in the page cache. decode_budget_q2_0 at one to four threads gave:
| threads | llama.cpp b11192 (min / median) | bankml (min / median) | ratio | matmul-only ceiling |
|---|---|---|---|---|
| 1 | 4.145 / 4.161 s | 0.435 / 0.439 s | 9.52Γ | 0.24 β 2.30 tok/s |
| 2 | 2.345 / 2.500 s | 0.284 / 0.313 s | 8.27Γ | 0.43 β 3.53 tok/s |
| 3 | 2.185 / 2.221 s | 0.231 / 0.232 s | 9.45Γ | 0.46 β 4.32 tok/s |
| 4 | 2.179 / 2.202 s | 0.223 / 0.228 s | 9.78Γ | 0.46 β 4.49 tok/s |
This reproduces the 0.0.3β0.0.6 figure (0.23β0.25 s at three threads, about 9Γ) on today's code, and it confirms the explanation above: the kernels never slowed down. The gates in between measured a laptop that could not keep the model in memory.
JSON mode's cost (0.3.3) β laptop
The grammar costs time only in three places, and the gate measures each:
- The check. The token the chain drew is run through the grammar alone: one candidate, microseconds.
- The mask. When that token breaks the grammar, the grammar runs over the whole vocabulary of 151,669 tokens, and the chain draws again.
- The accept. The chosen token advances the grammar's stacks.
oracle_grammar_masks times the mask on every one of its 1,645 recorded steps, over 12 grammars.
oracle_json_mode{,_ternary} time the grammar's whole share of real constrained answers: check, mask, accept and
prefill, divided over every generated token.
| measure (gate 0.3.3, laptop) | Bonsai-8B Q1_0 | Ternary-Bonsai-8B |
|---|---|---|
one whole-vocabulary mask, 151,669 tokens (oracle_grammar_masks: 1,645 masks over 12 grammars; the vocabulary is shared) |
median 24.2 ms Β· p90 47.0 Β· max 101.0 | the same |
| tokens redrawn under the mask, in the recorded answers | 152 of 860 (17.7 %) | 49 of 534 (9.2 %) |
| grammar time per generated token: check, mask when needed, accept, and the prefill, spread over the answer | 3.90 ms | 3.95 ms |
for scale: the matmul-only decode budget at 3 threads (decode_budget_*, same gate) |
0.389 s | 0.279 s |
So JSON mode costs about 1 % of a 1-bit step and 1.4 % of a ternary step on average. When a token has to be
redrawn, that step costs one mask more, about 6 % of a 1-bit step at the median. An answer the model writes as valid
JSON on its own costs almost nothing: bankml generate --json on the cat prompt redrew no token and spent 0.4 ms of
grammar per token.
The mask was a straight port of llama.cpp's reject_candidates, which walks every candidate's code points per grammar
stack. 0.3.9 replaced it with a trie over the vocabulary's code points, under the same oracle: see
the next section.
F16 and the Llama graph (0.3.4) β laptop, against llama-server b11192
The models O4 opened: SmolLM2-135M-Instruct and mindX's own mindx-gen39 (SmolLM2-135M + a LoRA, merged), both F16
with tied embeddings (270 MB), and Bonsai-1.7B Q1_0. Both engines at 3 threads, greedy, the same chat messages,
alternated in pairs (bankml generate against llama-server -t 3 -c 2048 -np 1 --jinja --reasoning off, its
/v1/chat/completions timings, cache_prompt: false). The machine was not idle (load average 1β2.8): read the ranges,
not single numbers.
| model, case | bankML | llama-server b11192 | ratio |
|---|---|---|---|
| SmolLM2-135M-Instruct F16, 45-token prompt, 128 tokens decoded | decode 38.0β38.9 tok/s (3 runs) | 40.4β42.7 tok/s | β 0.93Γ |
| the same, prompt | 0.3β0.4 s | 0.4 s (115β119 tok/s) | β 1Γ |
| SmolLM2-135M-Instruct F16, 858-token prompt (docs/OLLAMA.md's first 2,400 characters), 64 decoded | prompt 8.6β9.6 s, decode at ~900 cells 27.2β30.1 tok/s | prompt 7.8β8.6 s (100β110 tok/s), decode 20.2β29.9 tok/s | prompt β 0.9Γ, decode β 1Γ |
| Bonsai-1.7B Q1_0, 29-token prompt, 64 decoded | 8.4β8.6 tok/s (2 runs) | 10.6β10.9 tok/s | β 0.8Γ (the 1-bit decode gap of TODO 0.4.0, O8) |
mindx-gen39 has SmolLM2-135M's shapes, so it runs at SmolLM2's speed. Reproduce: bankml generate .models/SmolLM2-135M-Instruct-F16.gguf --max 128 < messages.json with BANKML_THREADS=3, beside a llama-server on the
same file.
What it took, each change the same bits (every oracle reran after each):
| change (0.3.4) | SmolLM2 decode, 45-token prompt | 858-token prompt |
|---|---|---|
| first correct build (scalar f16 helpers in attention, a linear scan per tensor lookup) | 12.8 tok/s | 18.5 s |
| the reference kernel's f16 dot, scale and accumulate on F16C + FMA; tensors by hash map; norms read once | 36.1 tok/s | β |
| tinyBLAS's 4 Γ 3 register blocking for F16 prompts, the activations widened once per product | β | 14.4 s |
| heads and rows of attention on the pool | 37.1 tok/s | 10.8 s |
| the tiled kernel widening each 64-cell tile once per 16 rows (compiled with FMA and F16C) | β | 7.8β9.6 s |
Rejected, with its numbers: a bounded spin before the pool's workers sleep (to save the ~15.6 Β΅s condvar wake per
matmul, bench_pool_overhead). On this 2-core SMT laptop the spinning threads took the siblings' cycles:
pool.run went from 15.6 Β΅s to 162.6 Β΅s per call, and decode did not move (33.7β36.1 tok/s). Reverted.
The bits did not move with any of it: F16 products are checked against the shipped library at every shape class
(oracle_ggml_b11192_f16), and the end-to-end oracles (oracle_llama_server_llama_f16, the live serve and JSON
oracles) compare whole answers with llama-server's.
The 8B models gained from the same attention work (the 0.3.3 and 0.3.4 gates on this laptop, wall time of the same oracle runs, token-identical in both; not a controlled benchmark β the machine's load differs between the runs):
| gate oracle (same requests, same tokens) | 0.3.3 | 0.3.4 |
|---|---|---|
oracle_greedy_llama_server_long (Bonsai-8B Q1_0, 6 prompts of 64+ tokens) |
435 s | 261 s |
oracle_greedy_llama_server_deep (Bonsai-8B Q1_0, 600 tokens past 256 cells) |
756 s | 320 s |
oracle_native_serve (Bonsai-8B Q1_0, 9 turns) |
400 s | 192 s |
oracle_json_mode (Bonsai-8B Q1_0, 23 answers) |
901 s | 541 s |
oracle_json_mode_ternary (Ternary-Bonsai-8B, 13 answers) |
454 s | 242 s |
decode_budget_q1_0, matmuls only, 3 threads (median) |
0.389 s | 0.397 s |
The matmul budget did not move, so the gain is the work around the matmuls: before 0.3.4 the reference attention kernel converted every f16 value in software, one at a time, and every tensor lookup scanned the list.
One paired check after the gate, end to end, Bonsai-8B Q1_0, the same 29-token prompt, 32 tokens, 3 threads each
(bench.sh-style pairs as above; load average 2.9β3.7, so read it as a hint, not a result): bankML decode
2.59 and 2.30 tok/s, llama-server 2.33 and 2.48 tok/s; prompts 8.7 / 10.5 s against 10.9 / 10.1 s. TODO 0.4.0
recorded 1.9β2.0 against 2.8 before. Whether 1-bit decode is now at parity needs the pinned, idle-machine
measurement (testing/pinned.sh); it is not claimed here. The method is in
1-bit decode against llama-server.
The grammar mask, through a trie (0.3.9, unreleased)
A whole-vocabulary mask now walks a trie of the vocabulary's code points, built once: each grammar stack meets a
shared prefix once (grammar.rs). oracle_grammar_masks, over the same 1,645 masks as the 0.3.3
table above (CHANGELOG 0.3.9):
| one whole-vocabulary mask, 151,669 tokens | trie (0.3.9) | the port of reject_candidates |
|---|---|---|
| median | 2.94 ms | 38.8 ms |
| p90 | 49.3 ms | 80.6 ms |
The median is 13Γ faster. The oracle computes every mask both ways: 196 / 196 runs and 1,645 / 1,645 masks are identical to llama.cpp b11192 by each. The port's median here (38.8 ms) is higher than the 0.3.3 gate's 24.2 ms. They are different runs on the same laptop; compare the two columns of one run, not runs with each other.
The q8_0 KV cache (0.3.9, unreleased): memory
BANKML_CACHE_TYPE=q8_0 keeps K and V as q8_0 blocks, as llama.cpp's --cache-type-k/v q8_0: 53 % of the f16
cache's bytes (q8_0 stores 34 bytes per 32 values, f16 64). Its answers are token-identical to llama-server so
configured (kv_oracle_live, 6 / 6; CHANGELOG 0.3.9). Its speed is not measured yet. A q8_0 cache has no tiled or
split-KV attention kernel, as in ggml, so long prefills are expected to be slower than with f16
(forward.rs).
1-bit decode against llama-server (0.3.9, pending)
The open question of 0.4.0 is whether 1-bit decode is at least at llama-server's speed. The paired checks above were
taken on a loaded laptop and are hints. The method for the answer is testing/decode_ab.py:
BANKML_GGML_LIB=<b11192 release dir> BANKML_PIN_CPUS=1,2,3 BANKML_PIN_MEM=3000M \
testing/pinned.sh python3 testing/decode_ab.py Bonsai-8B-Q1_0 [rounds] [threads]
- Each round starts a fresh llama-server b11192 and a fresh
bankml serve --native, alternating which goes first. - Both get the same greedy request with the prompt cache off, and each side's own
timingsare read (prompt_per_second,predicted_per_second). - The answers must be token-identical, or the round is refused.
- It prints every round and the medians; the ratio is bankML Γ· llama-server (above 1: bankML faster).
No result is recorded here yet. It will be, with the machine's load, when it is run on an idle machine.