Download docs/oracles.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 50.3 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/oracles.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/oracles.md
-
curl -L -o oracles.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/oracles.md
Oracles in bankml
What an oracle is here
An oracle is an answer bankml did not compute and cannot influence, against which its own answer is compared. bankml's rule since 0.0.1 is that a result counts only if an oracle has checked it: the same bits first, then the speed. A faster kernel that is not bit-identical to its oracle is not a result; a construction (a hash, a Merkle root, an ABI encoding) that does not reproduce a published value is not trusted.
The strongest oracle is the compiled reference itself: llama.cpp b11192's own shared libraries, called in-process
through their exported symbols on the same bytes bankml reads. The others are published values: test vectors from
standards (FIPS 180-4, RFC 6962), values published by the systems bankml interoperates with (Savante's doctrine root,
the THOT spec's vectors, Hugging Face's and Ollama's sha256 for a model file), and independent implementations
(pycryptodome, Foundry's cast, the Python guard bankml's Rust guard was ported from).
This file lists every oracle bankml uses, what it checks, how to run it, and what it last found. The results are in
testing/results/, one record per release, written by
testing/release_gate.sh. The latest record is 0.3.6's; figures marked 0.3.7, 0.3.8 or
0.3.9 are from CHANGELOG.md for versions not yet released, and reach testing/results/ when each
is cut. Each module's own oracles are also listed on its page in modules/, under How it is
verified.
The release gate, stage by stage
Every stage of testing/release_gate.sh, in the order it runs. A stage that fails stops the gate. The rows marked
regression compare bankml with itself or with its own scalar model; they guard against change, and by the rule at the
end of this file they are not oracles.
Always (no model needed)
| stage | checks | against | § |
|---|---|---|---|
cargo test, cargo test -p bankml-capi |
unit tests, the scalar models of §2, the gateway's receipts against a mock server | scalar models of ggml; a mock llama-server (regression where it is bankml against bankml) | §2, §5 |
cargo clippy -D warnings |
lint, both packages | — | — |
spdx_check.py |
every tracked source file names its licence in an SPDX header matching its layer | LICENSING.md | — |
test_gguf_guard.py |
the Python guard's own suite | its fixtures | §3 |
test_ui.py |
the UI's data layer: keccak256, THOT manifests, CIDv1, the RFC 6962 tree | published values | §4 |
test_console.py (0.3.7) |
the console's routes, loopback and CSP rules, the SELF block's units, the persona's doctrine root, the Infotags metadata | RFC 6962 roots; regression otherwise | §4 |
test_connectors.py |
PostgreSQL publish and load on a throwaway cluster; a tampered row refused | the manifest's bytes | §4 |
test_chain.py |
iNFT mint and load on a throwaway anvil devnet | the compiled contract | §4 |
test_models.py |
the model importer: the pin, the licence gate, resume, a carrier started and rolled back | published sha256 values | §4 |
guard_agree.py |
the Rust guard against the Python guard, JSON for JSON | the Python original | §3 |
capi_oracle.py --printf |
bankml_log byte for byte |
libc snprintf |
§5b |
With the b11192 release (BANKML_GGML_LIB), each a Rust test run by name (cargo test --release -- --ignored --exact), then the live checks against a running bankml serve --native:
| stage | checks | against | § |
|---|---|---|---|
schema_oracle.py, content_oracle.py (when LLAMA_SRC is set) |
re-record the schema grammars and the content rule | llama.cpp's own code in libllama-common.so |
§5e |
sort_oracle.cpp (when g++ is present) |
re-record std::sort's orders |
libstdc++ | §5f |
oracle_tokenizer |
token ids, 4,346 cases | llama-server /tokenize |
§1b |
oracle_chat_template |
prompts byte for byte, 317 conversations | llama-server /apply-template |
§1c |
oracle_forward_embed_norm, oracle_forward_qkv_rope, oracle_forward_attention, oracle_forward_attention_tiled, oracle_forward_attention_split, oracle_forward_swiglu_sweep |
layer 0 operation by operation, the three attention kernels, ggml's expf |
the shipped ggml | §1d |
oracle_forward_model, oracle_forward_model_ternary |
every layer, result_norm and the logits |
the shipped ggml graph | §1d |
oracle_greedy_llama_server, _ternary, _long, _deep |
greedy tokens, short, long and deep prompts | llama-server | §1d |
oracle_sample_llama_server |
seeded sampling, 40 continuations | llama-server | §1d |
oracle_native_serve |
three conversations, 9 turns: text, counts, cache reuse | llama-server /v1 |
§1d |
oracle_grammar_masks |
whole-vocabulary masks, both paths since 0.3.9 | libllama's grammar sampler | §5c, §5h |
oracle_json_mode, oracle_json_mode_ternary |
answers under JSON mode and user grammars | llama-server | §5c |
oracle_schema_grammars, oracle_json_schema, oracle_json_schema_ternary, oracle_json_schema_o4, oracle_json_content |
schema grammars, answers under schemas, the content rule | llama.cpp's own code; llama-server | §5e |
oracle_persona_layer |
mindX's persona Modelfile, two ways | llama-server with the persona as system message | §5e |
oracle_penalties, oracle_penalties_8b |
the repeat, frequency and presence penalties | llama-server | §5f |
oracle_samplers, oracle_samplers_8b (0.3.7) |
typical-p, top-n-σ, XTC, dynamic temperature, DRY | llama-server | §5f |
oracle_std_sort (0.3.7) |
std::sort's order of equal keys |
libstdc++ | §5f |
oracle_ggml_b11192_q8_0_kv_kernels (0.3.9) |
the q8_0 KV cache's quantizer and dot product |
the shipped ggml | §5h |
gpu_q1_0_mat_vec_bit_exact |
the GPU kernels on every usable card | the CPU kernel, itself proven against ggml | §1d (0.2.13) |
oracle_train_script, oracle_train_imprint |
mindXtrain's author and score stages | mindXtrain's Python | §1d (0.2.13) |
oracle_ggml_b11192_real_bonsai_1_7b, _real_bonsai_8b_q1_0, _real_ternary_bonsai_8b |
every weight, q8_0 rows, dot products |
the shipped ggml | §1 |
oracle_ggml_b11192_f16 |
F16 products, both of ggml's paths | the shipped ggml | §5d |
oracle_tokenizer_smollm, oracle_chat_template_chatml, oracle_forward_model_bonsai_1_7b, oracle_forward_model_llama_f16, oracle_llama_server_bonsai_1_7b, oracle_llama_server_llama_f16, oracle_native_serve_o4 |
the O4 models: tokenizer, templates, whole model, llama-server's tokens, conversations and JSON mode | llama-server; the shipped ggml | §5d |
ab_vs_ggml, ab_vs_ggml_q2_0, bench_q1_0_prefill_act, bench_memory_floor, decode_budget_q1_0, decode_budget_q2_0 |
speed: kernel A/Bs (each first checks agreement), the memory floor, one token's matmul budget | the shipped ggml's own kernels; measurements, not oracles | §1, PERFORMANCE.md |
serve_oracle_ollama_shape |
the conversations through /api/chat, /v1, and /v1 after a reload |
llama-server's record | §1d (0.3.1) |
o4_live |
the conversations, JSON mode and schemas live, on Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 |
llama-server's records | §5d, §5e |
persona_oracle_live |
the created mindx-gen39 through /api/chat, /api/generate, /v1; /api/ps |
llama-server's record | §5e |
penalty_oracle_live, sampler_oracle_live |
the penalties and the sampler chain through /v1 and /api/chat |
llama-server's records | §5f |
context_oracle_live (0.3.8) |
the context limit | llama-server's record | §5g |
slot_oracle_live (0.3.8) |
slot save, restore and erase | the engine's own empty-slot answer; llama-server's refusals | §5g |
session_oracle_live (0.3.8) |
interleaved conversations; simultaneous requests | llama-server's record; the queue against itself | §5g |
kv_oracle_live (0.3.9) |
answers over a q8_0 KV cache |
llama-server with --cache-type-k/v q8_0 |
§5h |
logprobs_oracle_live (0.3.8) |
logprobs, streamed and not | llama-server's record | §5g |
json_oracle_live, json_schema_oracle_live |
JSON mode and schemas live on the 8B model | llama-server's records | §5c, §5e |
capi_chat_oracle |
bankml_chat |
serve --native; llama-server's record |
§5b |
When their inputs are present
| stage | checks | against | § |
|---|---|---|---|
oracle_convert_b11192 (per directory under .models/convert/) |
bankml convert's GGUF, sha256 for sha256 |
b11192's convert_hf_to_gguf.py |
§5e |
oracle_name_heuristics (when BANKML_LLAMA_SRC is set) |
model-name heuristics, 168 ids | gguf-py | §5e |
1. The ggml oracle: kernels bit-exact against the compiled llama.cpp
What it checks
For each real model, three things, every one to the bit:
| # | quantity | bankml side | ggml side (llama.cpp b11192, shipped binaries) |
|---|---|---|---|
| 1 | every weight of every low-bit tensor, dequantized to f32 | q1_0::dequantize_row, q2_0::dequantize_row |
dequantize_row_q1_0 / dequantize_row_q2_0 in libggml-base.so |
| 2 | activation rows quantized to q8_0 (32 signed bytes and a half-precision scale per block) |
quantize_row_q8_0 |
quantize_row_q8_0 in libggml-cpu-haswell.so |
| 3 | dot products of a weight row with a q8_0 row |
vec_dot_ref (scalar model), vec_dot / vec_dot_act (AVX2), vec_dot_act_scalar |
ggml_vec_dot_q1_0_q8_0 (AVX2 path), ggml_vec_dot_q1_0_q8_0_generic, ggml_vec_dot_q2_0_q8_0 |
Dequantized tensors are compared by the sha256 of their little-endian f32 output (a whole 8-billion-weight model,
tensor by tensor, without holding it in memory); quantized rows byte for byte; dot products by their f32 bit
patterns (to_bits()), not within a tolerance.
Why "the compiled library" and not "the source"
The first port of ggml's generic C, done faithfully from the source, disagreed with the shipped library in the last
bit. Reading the disassembly showed why: the compiler had fused a multiply and an add into one FMA instruction, which
rounds once instead of twice. A source-level port cannot see that; the binary oracle catches it. bankml's kernels
reproduce the float order and the fused multiply-adds of the shipped libggml-cpu-haswell.so (the variant ggml's
backend scorer loads on an AVX2 CPU without AVX-512), and the oracle proves it on real tensors.
It also shows the oracle can tell float orders apart. For the ternary format ggml ships two builds of the same generic
C: the haswell variant (FMA) and the baseline x64 variant (SSE, no FMA). They disagree with each other on 19 of the
762 recorded ternary dot products. bankml carries a model of each (vec_dot_ref for haswell, vec_dot_ref_nofma for
x64) and matches each one in 762 of 762.
How it runs
Record (once per model):
testing/ggml_oracle.pyloads the release'slibggml-base.soandlibggml-cpu-haswell.so(and, for Q2_0,libggml-cpu-x64.so) withctypes, callsggml_cpu_init(), memory-maps the GGUF and writes:dequant.tsv: tensor name · element count · sha256 of ggml's f32 output, for every tensor of the type;vecdot.bin: for sampled rows of every tensor (3 per tensor by default): the f32 activationx, ggml'sq8_0of it, and ggml's dot products (AVX2 and generic).
The release tarball is checked by sha256 before use (
llama-b11192-bin-ubuntu-x64.tar.gz,34cf6fa5…81ec7).Compare (every gate): the Rust tests
oracle_ggml_b11192_real_*(inq1_0.rsandq2_0.rs,#[ignore]d because they need the models) re-derive every recorded quantity from the same file through bankml's own memory map and assert equality.A/B against the live library (every gate):
ab_vs_ggmlandab_vs_ggml_q2_0dlopenthe shipped haswell library directly (no crate,dlopen/dlsymdeclared by hand) and time ggml's kernel and bankml's on the same real weights in one process, after first checking they agree.
What it found (0.1.7–0.2.0 gates; identical in each, and again under Rust 1.99 in 0.3.2)
The gate runs these as oracle_ggml_b11192_real_bonsai_1_7b, oracle_ggml_b11192_real_bonsai_8b_q1_0 and
oracle_ggml_b11192_real_ternary_bonsai_8b.
| model | tensors | weights | q8_0 rows | dot products |
|---|---|---|---|---|
Bonsai-1.7B, Q1_0 |
197 | 1,719,904,256 | 788 byte-exact | 788 bit-exact (AVX2 and vec_dot_act); generic port == ggml generic 788/788 |
Bonsai-8B, Q1_0 |
254 | 8,188,239,872 | 762 byte-exact | 762 bit-exact; generic 762/762 |
Ternary-Bonsai-8B, Q2_0_g64 |
254 | 8,188,239,872 | 762 byte-exact | 762 bit-exact vs haswell (reference, AVX2 and scalar paths); no-FMA model == x64 build 762/762 |
Run it
# once: the oracle files (numpy + the b11192 release)
python3 testing/ggml_oracle.py .models/Bonsai-1.7B-Q1_0.gguf /path/to/llama-b11192 .models/oracle
python3 testing/ggml_oracle.py .models/Bonsai-8B-Q1_0.gguf /path/to/llama-b11192 .models/oracle-8b-q1
python3 testing/ggml_oracle.py .models/Ternary-Bonsai-8B-Q2_0_g64.gguf /path/to/llama-b11192 .models/oracle-ternary
# every time
cargo test --release -- --ignored oracle_ggml_b11192 --nocapture --test-threads=1
BANKML_GGML_LIB=/path/to/llama-b11192 cargo test --release -- --ignored ab_vs_ggml --nocapture --test-threads=1
1b. The tokenizer oracle: token-identical to llama.cpp (0.2.1)
P3, bankml's own forward pass, starts with the tokenizer. testing/tokenizer_oracle.py
asks a running llama-server (b11192, the Bonsai / Qwen3 vocabulary) to tokenize a corpus, with special tokens parsed
and not, and records every answer. The corpus is:
- every document in the repository and Savante's canon texts;
- the chat template's markers;
- hand-picked edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with joiners, mathematical alphanumerics;
- a seeded fuzz set of 2,000 strings drawn from 17 Unicode ranges.
tokenizer::tests::oracle_tokenizer re-derives every case with bankml's tokenizer (tokenizer.rs, no crates) and
requires the same ids in the same order. Last result: 4,346 of 4,346 cases token-identical (the corpus grew at 0.3.4; 4,258 at 0.2.1), on the Qwen2 and the SmolLM2 pre-tokenizer each.
Two details the oracle settles:
- Which special tokens split the text. Qwen3's
<think>markers are USER_DEFINED and split the text even when special tokens are not parsed; CONTROL tokens such as<|im_start|>do not. - The pre-tokenizer's letter class. It is Unicode general category L, which is not
char::is_alphabetic. It is generated intounicode_letters.rs.
1c. The chat-template oracle: byte-identical prompts (0.2.2)
A conversation becomes a prompt through the model's chat template, a Jinja program inside the GGUF. bankml does not
run Jinja. chat.rs writes the Bonsai / Qwen3 template's rules out, and check_template accepts only that template,
identified by the sha256 of its text. testing/template_oracle.py records
llama-server's own /apply-template for 317 conversations. They cover:
- a system prompt first, later, or absent;
- assistant turns with and without
<think>blocks, before and after the last real user query; reasoning_content;- runs of tool results;
- user messages that look like tool responses;
- special markers and Unicode inside content;
- 300 random conversations.
oracle_chat_template requires every prompt byte-identical: 317 of 317. The oracle found one server behaviour
that the template alone would not predict: an empty reasoning_content is dropped before templating.
1d. The forward-pass oracle: llama.cpp's own operations, from the shipped ggml (0.2.3)
The release has no tool that prints a model's intermediate values, so the oracle builds them the way the kernel oracle
does. testing/forward_oracle.py drives the shipped libggml through ctypes with
the same operations llama.cpp's Qwen3 graph uses and computes them with the release's own CPU backend. For 300 real
token ids (the chat markers, the table's edges, and tokens of the tokenizer's corpus) it records the sha256 of each
row's f32 bytes:
get_rowson the Q1_0 token embedding: llama.cpp'sinp_embd;rms_normwith the model's epsilon, thenmulbyblk.0.attn_norm.weight: llama.cpp'sattn_norm-0.
bankml's forward.rs computes the same in the float order read from b11192's ops.cpp:
- the sum of squares accumulated in double, one float product at a time, in order;
- the mean rounded to float;
scale = 1 / sqrtf(mean + eps);- each output
(x · scale) · w, with no FMA.
oracle_forward_embed_norm requires 300 of 300 rows bit-exact for both.
Step four (0.2.4): Q, K, V, the head norms and RoPE. The oracle also builds layer 0's attention inputs on a 28-token chat prompt:
mul_matof the Q1_0attn_q,attn_kandattn_vweights with the normed rows;reshapeto 128-wide heads,rms_normandmulbyattn_q_normandattn_k_norm;rope_extin NEOX mode with the YaRN parameters llama.cpp's context derives (freq_scale0.25,ext_factor1,attn_factor1.0, whose bits the record carries, beta 32 and 1,n_ctx_orig16,384, base 10⁶).
RoPE runs twice: at positions 0–27, and at positions 7 to 63,214, where theta is a long product.
oracle_forward_qkv_rope requires 140 of 140 rows bit-exact.
This step showed why the oracle is the shipped binary and not the source. Written as ops.cpp reads, bankml's RoPE
matched only 30 of 84 rows, while the projections and norms before it were exact. The disassembly of
libggml-cpu-haswell.so shows GCC's FMA contraction of the rotation, the YaRN mix and the magnitude term. With those
three written as the same mul_adds, every row matches.
Step five (0.2.5): attention. The oracle stores K and V through cpy to f16, as llama.cpp's cache does. It
builds the causal f16 mask and runs flash_attn_ext, which is llama-server's default attention on the CPU, with F32
accumulation set as llama-graph sets it. The output then goes through wo (mul_mat) and the residual add.
oracle_forward_attention requires 84 of 84 rows bit-exact (kqv_out, attn_out, ffn_inp, 28 tokens).
The reproduction follows ggml's reference path:
- the f16 dot's four 8-lane FMA accumulators and their reduction tree;
- an online softmax whose V accumulator is rounded to f16 at every step.
The oracle rejects a near miss: with the softmax sum written as an FMA, only 59 of 84 rows match. At 0.2.5 two other kernels were not yet covered; each got its own oracle in 0.2.9 and 0.2.10 (below):
- the tiled kernel, for 64 or more query rows;
- the split-KV kernel, for a decode whose padded KV length reaches 512 (llama.cpp pads to multiples of 256, so from 257 cells in use).
Step six (0.2.6): the feed-forward block; layer 0 complete. The oracle continues from ffn_inp:
rms_normandmulbyffn_norm;mul_matbyffn_gateandffn_up;swiglu_split;mul_matbyffn_down;add, givingl_out, the layer's output.
oracle_forward_attention requires 112 of 112 rows through l_out. SiLU in the shipped build uses ggml's own
vectorized expf, a polynomial in FMAs, not libm. bankml reproduces it lane by lane.
A second check, oracle_forward_swiglu_sweep, feeds 24,600 values across ±120 and the edges straight to
ggml_swiglu_split and requires every output's bits (24,600 of 24,600). The sweep exists because the layer check
alone was not enough: with libm's expf, l_out still matched 27 of 28 rows, since q8_0 quantization absorbs most
of the difference, while the sweep matched only 19,245 of 24,600.
Step seven (0.2.7): the whole model, and llama-server itself.
testing/model_oracle.pycomputes the whole Qwen3 graph in the shipped ggml: every layer,output_normandmul_matbyoutput.weight. It runs one layer per ggml context and carries the residual stream between contexts as f32 bytes, which is the same arithmetic as one graph in a fraction of the memory.oracle_forward_modelrequires 1,064 of 1,064 rows bit-exact: 36 × 28l_outrows, 28result_normrows, and 28 logit rows of 151,669 values each.testing/greedy_oracle.pysteps outside the graph. It asks the running llama-server to render 6 chat prompts (/apply-template), tokenize them (/tokenize) and continue them greedily (/completion, top-k 1, prompt cache off).oracle_greedy_llama_serverrequires bankml's own forward pass to produce the same tokens on every prompt (6 of 6, 164 tokens), and to end the turn where the server did.
Both stayed inside the range bankml reproduced at 0.2.7: prompts under 64 tokens, contexts up to 256 cells (the largest case uses 80). Steps nine and ten removed both limits. llama.cpp pads the KV length to multiples of 256, so a decode beyond 256 cells takes the split-KV kernel.
Step eight (0.2.8): the ternary model. Both whole-model checks run again on Ternary-Bonsai-8B (Q2_0_g64), whose every matrix, including the embedding and the output, is ternary:
oracle_forward_model_ternarychecks bankml against the shipped ggml graph: 1,064 of 1,064 rows.oracle_greedy_llama_server_ternarychecks against llama-server running the ternary model, with Savante's flags, on a spare port: 6 of 6 prompts, 140 tokens.
The greedy comparison now covers exactly the tokens the server generated. Its list includes the end-of-turn token when it produced one. Where it stopped otherwise, that is its stopping policy, not a token choice.
Step nine (0.2.9): long prompts and ggml's tiled kernel. llama.cpp computes a prompt in micro-batches of up to
512 tokens. A micro-batch of 64 rows or more takes flash_attn_ext_tiled, a different algorithm from the reference
path: f32 Q, a SIMD GEMM per 64-cell KV tile, a vectorized softmax summed in double, and an f32 accumulator.
- The forward oracle runs layer 0 on a 150-row micro-batch in the shipped ggml.
oracle_forward_attention_tiledrequires 150 of 150 rows bit-exact. The reference kernel would match only 1, the first row, so the kernel choice is itself checked. greedy_oracle.py --longrecords llama-server's continuations for 6 prompts of 111–116 tokens.oracle_greedy_llama_server_longrequires the same tokens on all 6.
Step ten (0.2.10): long contexts and ggml's split-KV kernel. A single-token decode over 512 or more padded cells cuts the cells into one chunk per llama.cpp thread and reduces the partial softmaxes, so the thread count shapes the bits.
- The forward oracle records one decode row at 7 context lengths (257–1,000 cells) and at 3 and 4 threads.
oracle_forward_attention_splitrequires 14 of 14. greedy_oracle.py --deeprecords 3 continuations of 200 tokens that run to about 300 cells.oracle_greedy_llama_server_deeprequires the same 600 tokens.
The 4-thread cases did double duty. Besides showing the dependence on the thread count, they exposed the reduction's FMA contraction: at 3 threads the first chunk always held the maximum, where the fused and plain forms agree.
Step eleven (0.2.11): sampling. testing/sample_oracle.py has llama-server sample 40 continuations with fixed
seeds, and keeps the parameters the server reports it ran (generation_settings). oracle_sample_llama_server
replays them through bankml's forward pass and sampler.rs and requires every token (40 of 40, 1,175 tokens).
The cases range over temperature 0–1.5, top-k 5–128, top-p and min-p. Top-k is libstdc++'s partial_sort, ported
line for line, because the order it leaves tied logits in decides which token a draw lands on.
0.2.13: the GPU and mindXtrain.
gpu_q1_0_mat_vec_bit_exactruns both Q1_0 GPU kernels on every usable card against the CPU kernel, which is bit-exact against ggml: on the Vega 3, every row of every shape from 64×128 to 12288×4096.bankml gpu --verifyruns the same check for a user, and bankml trusts a card only after it passes.oracle_train_scriptchecks bankml's author stage against mindXtrain'sscripts.pyon 84 scripts, byte for byte.oracle_train_imprintchecks the score stage againstscore_imprinton 3,000 randomized cases, value for value.
0.2.14: the GPU inside the forward pass. With a verified card present, Weights::open gives it a share of every
1-bit matrix's rows, so every token oracle runs with the GPU working. The whole-model check passes at 1,064 of
1,064 rows with the Vega 3 on 26 % of every matrix.
The on-card oracle learned the lesson of this release. It now includes layer-shaped data: scales and magnitudes
that vary per block, so products are inexact in f32. On that data the card's unfused Fma differed by one ulp,
where random data had let it pass. bankml now computes the FMA exactly (Boldo–Melquiond), and with the driver's
Fma put back, bankml gpu --verify refuses the card.
0.3.0: whole conversations through the server. testing/serve_oracle.py holds three Savante-style
conversations, three turns each, with fresh seeds at temperature 0.3. It sends them to a fresh llama-server's
/v1/chat/completions, the path Savante uses, so the server's prompt cache carries each conversation from turn to
turn, and records every answer with its prompt count, completion count and the prompt tokens taken from the cache.
oracle_native_serve replays the conversations through bankML's native engine in one process and requires all of
it on 9 of 9 turns: the same text, the same counts, the same cache reuse. The rule it reproduces is visible in
the numbers. Turn 2 reuses 49 of turn 1's 53 prompt tokens, because the empty <think> block that ended turn 1's
prompt is not rendered when that answer becomes history. A new conversation reuses the 38 tokens of the shared
system prompt.
0.3.1: the Ollama shape. testing/serve_oracle.py --bankml starts bankml serve --native with the 8B 1-bit model
on spare ports and sends the same conversations three ways, each from an empty slot (the model unloaded, then loaded
again, so each run starts as a fresh llama-server does): through Ollama's /api/chat, with the first conversation
streamed as NDJSON and the options mapped from Ollama's names; through /v1/chat/completions; and through /v1 once
more after an unload in which /v1 itself loads, and verifies, the startup model. Every turn must give the same text,
prompt and completion counts and cache reuse on every path, equal to llama-server b11192's record. So the Ollama
path inherits the conversation oracle, and a reload changes nothing. The gate runs it as serve_oracle_ollama_shape.
Every step of the forward pass above was added to these oracles before it counted, and every later feature (§5b–§5h) follows the same rule.
2. Scalar models: every fast path against its own reference
Between the real-model oracle runs, the kernels are held to a scalar model of ggml, on synthetic inputs, in the
ordinary cargo test:
| check | where | cases |
|---|---|---|
| half-precision conversion, all 65,536 values, and against the CPU's F16C instruction | q1_0.rs |
65,536 |
q8_0 rounding (round-half-to-even, scale = 127 / max|x|) |
q1_0.rs |
edge cases |
AVX2 Q1_0 dot == scalar model of ggml |
q1_0.rs |
3,500 |
AVX2 and portable Q2_0 dot == ggml model |
q2_0.rs |
3,300 |
| matrix–vector, matrix–matrix and threaded paths == per-pair dots | q1_0.rs, q2_0.rs, par.rs |
every shape tested |
| the 0.0.4 prefill tile == per-pair | q1_0.rs (act_tile_bit_exact) |
every tile |
| the generic port == f64 arithmetic on dequantized data (the mathematical definition) | q1_0.rs |
random blocks |
A speed-up is admitted only after these and §1 pass on the same build. Kernels that were faster but not exact, or
exact but not faster beyond noise, are kept with their numbers in testing/experiments/
(0.0.5).
3. The guard: Rust against the Python it was ported from
bankml's header guard (gguf.rs: play, refuse or need more bytes; the three low-bit traps; hostile headers) was
ported from minaiml's Python gguf_guard.py, vendored in testing/gguf_guard.py with its
suite. testing/guard_agree.py runs both on every synthetic case and on every real
model present and requires identical JSON. Last result: 28/28 agree.
4. Cryptographic constructions against published values
| construction | oracle | where |
|---|---|---|
| SHA-256 (the model pin) | FIPS 180-4 test vectors; since 0.1.8 the SHA-NI path must also equal the portable rounds on every length 0–1,000 (and 4 KiB, 64 KiB, split updates), and a real 1.16 GB model must hash to coreutils sha256sum's value and its published pin |
sha256.rs (fips_vectors, hardware_path_equals_portable_on_every_length) |
| the model pin | the sha256 the publisher lists: the FORK.json of the PYTHAI fork, a Hugging Face repository's LFS sha256 at a fixed revision, or the Ollama registry's layer digest; every import is hashed as it streams and kept only if equal | bankml.rs (pin), sAGI/models.py |
| keccak256 (pure Python) | pycryptodome's keccak on every input length 0–400 bytes; and Savante's published doctrine root 0x92fe83eb…ae137d0, reproduced from her persona |
sAGI/agents.py, testing/test_ui.py |
THOT manifests (sagi.thot_manifest/1) |
the spec's own test vectors (THOT_MANIFEST.md §5: savante@1fcca89, jaimla@8b57ccf, luvai@0c1eef7): bundle root, Merkle root, identity CID |
sAGI/thot.py, testing/test_ui.py |
| CIDv1 (raw, sha2-256, base32) | the CID of "abc" that mindX's rage.py and Savante's ledger construction give |
testing/test_ui.py |
.history Merkle tree (RFC 6962) |
the Certificate Transparency reference roots for 1, 3 and 8 leaves | testing/test_ui.py |
| iNFT ABI encoding | the compiled iNFT_7857 contract itself, deployed from its artifact on a throwaway anvil devnet: it must accept the calldata bankml encodes (simulate and mint), refuse what it should (missing role, a repeated content root), and read back exactly the values encoded |
sAGI/chain.py, testing/test_chain.py |
| Savante's canon | her own offline verifier bind/savante_verify.py: 12 of 12 commitments, doctrine root |
the Verifier tab |
5. The gateway: receipts against the text
bankml serve's receipts are checked end to end against a mock llama-server whose answers are known
(testing/cli.rs): the response_sha256 equals the sha256 of the text the mock sent (streamed and
not), request_sha256 equals the sha256 of the request body, and a model file changed after verification produces a
503 and no receipt. In use, the UI recomputes every answer's sha256 and marks ✓ or ≠ received!.
5b. The C API (0.3.2): the library against libc, and against the server
Two oracles hold the C API (capi/, CAPI.md), both run from C programs compiled with the system cc.
bankml_logagainst libcsnprintf(testing/capi/printf_oracle.c).bankml_logis a C-variadic function defined in Rust, and its formatter is bankML's own, so the oracle is the C library'ssnprintf. Each case passes the same format and arguments to both. The message delivered to the sink must equal snprintf's output byte for byte, length included. The cases cover:- integers at
INT_MIN/LLONG_MIN/SIZE_MAX, in every length; %fties (%.0fof 0.5, 1.5 and 2.5),%.1100fof the smallest subnormal,DBL_MAX, andinf/nanwith signs;- strings, NULL, and an unterminated buffer bounded by a precision;
%cof 0 and 255,%p,%%, widths, precisions and*.
Then the cases it does not support must give their exact markers. No argument may be read that cannot be typed, and
%nmust not write. The same comparison runs incargo test -p bankml-capi, from Rust.- integers at
bankml_chatagainstserve --nativeand llama-server (testing/capi/capi_oracle.py --chat, the gate'scapi_chat_oracle). The conversation oracle of §1d, carried across the C boundary:- on Ternary-Bonsai-8B, the same request bytes go to a live
serve --nativeand then tobankml_chatin a fresh process. They must give the same text, the same streamed pieces, the same counts and cache reuse, and the same receipt hashes; - on Bonsai-8B Q1_0,
bankml_chatmust give llama-server b11192's recorded turns; - Bonsai-1.7B and an unpinned file must be refused, with
serve's reasons.
- on Ternary-Bonsai-8B, the same request bytes go to a live
5c. JSON mode and grammars (0.3.3): llama.cpp's grammar sampler, and llama-server's answers
Two oracles hold O6's first cut (bankML/grammar.rs). Neither compares bankML with itself.
- The grammar engine against llama.cpp's own (
testing/grammar_oracle.py→oracle_grammar_masks).testing/grammar_oracle.cppdrives libllama b11192's public API (llama_sampler_init_grammar,_apply,_accept) on the pinned vocabulary, loaded vocab-only. The grammar functions inside llama.cpp are not exported, but the sampler that wraps them is, so the oracle is the code the server runs, not a model of it.The vocabulary as the grammar reads it. Every token's piece (
llama_token_to_piece, special tokens rendered) and whether it ends generation. bankML must agree on 151,669 of 151,669 tokens and on the end set of 6. That set includes<|fim_pad|>,<|repo_name|>,<|file_sep|>and token 128247</s>, which bankML's engine had not counted as ends before.Masks. The whole-vocabulary mask, before every token and after the last, is recorded as how many tokens are allowed and an FNV-1a hash of which. The runs span 12 grammars:
- llama-server's JSON-mode grammar, with its generation-prompt prefill;
- all 8 grammars in llama.cpp's
grammars/(json, json_arr, arithmetic, c, chess, english, japanese, list); - three written for what those miss: token terminals
<think>,!<|im_end|>and<[151644]>;.;{m,n}edges; comments and CRLF; astral and CJK ranges.
The inputs are valid, invalid, partial, escaped and unicode JSON, whitespace past
space's 20-character and 2-newline limits, prose, fences and truncations. Each goes in twice: tokenized as the tokenizer does, and as one token per byte, which splits every multi-byte character across tokens and exercises the partial-UTF-8 path. The result: 196 of 196 runs; 1,645 of 1,645 masks and 116 of 116 rejection points identical.
- llama-server's answers under a grammar (
testing/json_oracle.py --record→oracle_json_mode,oracle_json_mode_ternary). llama-server b11192 runs with Savante's flags and--verbose, so that each answer reports its tokens and the grammar it ran.- What is sent.
/v1/chat/completionswithresponse_format: {"type": "json_object"}on 7 prompts, greedy and seeded at temperatures 0.7, 1.0 and 1.3 (top-k 40). Some prompts invite JSON. Others tempt prose, a haiku or code, so that the grammar has to reject what the model wants: 152 of 860 tokens on the 1-bit model were redrawn under the mask. One is a nested object of 159 tokens, one is unicode, and two answers are cut bymax_tokens. The 1-bit model also gets two usergrammarrequests. - What is required. Each case runs from an empty cache. bankML's engine must give the same token ids, the end
token included (one answer ends on
<|file_sep|>), and the same raw text,content, finish reason and counts. The grammar text and generation prompt the server reports must be bankML's constant and prefill. - The result: Bonsai-8B Q1_0 23 of 23 (860 tokens); Ternary-Bonsai-8B 13 of 13 (534 tokens, 49 redrawn).
- The live check (
testing/json_oracle.py --bankml, the gate'sjson_oracle_live) sends a subset through a runningserve --native, each from an empty slot:/v1withresponse_format, streamed once;/v1withgrammar; and Ollama's/api/chatwithformat: "json". Every answer must equal the record: 8 of 8.
- What is sent.
5d. O4 (0.3.4): F16 products, the Llama graph, tied embeddings — the same oracles, three new models
Every oracle family the 8B models have, extended to Bonsai-1.7B (tied embeddings), SmolLM2-135M-Instruct and
mindx-gen39 (the Llama graph in F16), each against llama-server b11192 running that very file; plus one new kernel
oracle.
- F16 products against the shipped library (
testing/f16_oracle.py→oracle_ggml_b11192_f16). ggml multiplies an F16 weight two ways:ggml_vec_dot_f16for one column, llamafile's tinyBLAS for two or more (when the weight has a multiple of 4 rows and the row a multiple of 8). The oracle buildsggml_mul_matgraphs through ctypes and has the shippedlibggml-cpu-haswell.socompute them: SmolLM2's real matrices (layers 0 and 29, and the 49,152-row token table) by 1, 2, 3, 5 and 28 columns, and synthetic shapes that reach the tails (k % 32 ≠ 0) and the fallbacks (rows % 4 ≠ 0,k % 8 ≠ 0). Every F16 tensor widened byggml_fp16_to_fp32_rowis hashed too. The result: 211 of 211 tensors; 552,268 of 552,268 elements in 87 products (61 tinyBLAS, 26vec_dot). The two paths give different bits for the same row and column (the_two_reductions_differ), so the oracle checks the choice as well. - The whole model (
testing/model_oracle.py, now for tied embeddings, F16 and the Llama graph:rope_extmode 0, no Q/K norms,ext_factor0): every layer'sl_out,result_normand the logits. An F16 model is replayed as one micro-batch with every row output, because that is what the oracle's graph computes. Bonsai-1.7B 840 of 840 rows, SmolLM2-135M-Instruct and mindx-gen39 800 of 800 each (oracle_forward_model_bonsai_1_7b,oracle_forward_model_llama_f16). - The tokenizer (
tokenizer_oracle.py 127.0.0.1:PORT .models/oracle-tokenizer-smollm→oracle_tokenizer_smollm): the same corpus on SmolLM2's vocabulary, 4,346 of 4,346, with two new edge strings: bytes SmolLM2 has no token for, and plane-4 characters. Re-recorded on the Qwen3 vocabulary, the grown corpus caught a bug:</s>(NORMAL in the file, CONTROL in llama.cpp by its name) was taken as text in 2 of 4,346 cases. Fixed; 4,346 of 4,346. - The templates (
template_oracle.py, written per model ascases-<stem>.jsonl→oracle_chat_template_chatml): SmolLM2-Instruct's and mindx-gen39's ChatML, 317 of 317 each. - llama-server's tokens (greedy short,
--long,--deep; seeded; conversations; JSON mode, each recorded from that model's server;oracle_llama_server_bonsai_1_7b,oracle_llama_server_llama_f16,oracle_native_serve_o4): SmolLM2 6, 6, 3 and 40 of 40; mindx-gen39 6, 6, 3 and 40 of 40; Bonsai-1.7B 6 and 40 of 40; conversations 9 of 9 and JSON mode 23 of 23 on each. On a ChatML template llama-server builds a different JSON grammar (no<think>in its root); the record shows it, and bankML carries it per template. - Live and through the C API:
serve_oracle.py --bankml STEM [NAME]andjson_oracle.py --bankml STEM [NAME]on each model (mindx-gen39 asked for asmindx-gen39), andcapi_oracle.py --chatagainst each model's record. The gate runs the live pair, withjson_schema_oracle.py --bankmlsince 0.3.5, as itso4_livestage.
5e. JSON schemas and bankml create (0.3.5): llama.cpp's converter, parser and converter script, and the persona
- The schema grammar (
testing/schema_oracle.{cpp,py}→oracle_schema_grammars): llama.cpp b11192's ownjson_schema_to_grammarandcommon_chat_templates_apply, called inside the release'slibllama-common.so(no model), on 173 schemas — llama.cpp's 81 test cases, Pydantic-shaped schemas like mindX's, edge cases and every refusal path — with the template read from each model that carries it (Bonsai's Qwen3, SmolLM2-Instruct's ChatML, mindx-gen39's ChatML). 148 grammars byte-identical bare and on the chat path, on each template; every refusal with llama.cpp's message (24, 20). The ChatML wrapping (no reasoning rules, another root) was read from this output. - The answers (
testing/json_schema_oracle.py --record STEM→oracle_json_schema,oracle_json_schema_ternary,oracle_json_schema_o4): llama-server b11192 asked under schemas through all three request shapes, greedy and seeded, each from an empty cache, with answers cut bymax_tokensinside an object, a top-level string and a number. Required: the same token ids (end token included), raw text, content, finish and counts, and the server's grammar and generation prompt are bankML's. Bonsai-8B 28 of 28, Ternary-Bonsai-8B 11 of 11, Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 56 of 56 each. Live (--bankml STEM [NAME], the gate'sjson_schema_oracle_liveando4_live):/v1(one streamed) and/api/chatformat: <schema>. - The content rule (
testing/content_oracle.{cpp,py}→oracle_json_content): llama.cpp'scommon_chat_parseitself, with the parser llama-server builds for each request, on every prefix of every recorded constrained answer and on edge cases: 30,063 of 30,063 texts on each template (36 texts with a reasoning block, which no grammar admits after its prefill, are refused by llama.cpp's parser and listed). It found two differences the answer oracles had not: unfinished escapes, and llama-server's raw-text answer for an empty parse; both fixed. - The converter (
testing/convert_oracle.py,oracle_convert_b11192,oracle_name_heuristics): bankml convert against llama.cpp b11192'sconvert_hf_to_gguf.py --outtype f16on the same directory, sha256 for sha256 — gen396b64c748…and SmolLM2ec30a679…, run directly in 0.3.5 (torch 2.14.1+cpu) — and gguf-py's name heuristics on 168 of 168 ids. - mindX's persona, end to end (
testing/persona_oracle.py→oracle_persona_layer, and live aspersona_oracle_live): promote.py's persona Modelfile created two ways (FROMthe merged directory;FROM mindx-gen39in place), the user's turns alone, against llama-server given the persona as the system message on the same GGUF: 27 of 27 token-identical for each; live through/api/chat,/api/generateand/v1from an empty slot each, and/api/psnames the derived model. (Without the empty slot, seeded answers moved: a cached prefix changes F16's product paths, so the live check starts each answer as the record did.)
5f. The penalties (0.3.6) and the rest of the sampler chain (0.3.7, unreleased)
- The penalties (
testing/penalty_oracle.py→oracle_penalties,oracle_penalties_8b). llama-server b11192 runs with Savante's flags from an empty cache, greedy and seeded, on 17 variants × 4 prompts made to repeat; the window reaches into the prompt, as llama-server fills it. Each case keeps the parameters the server read back. bankML must give the same tokens, and refuse what the server refuses with its message. mindx-gen39 56 of 56 (2,478 tokens), Bonsai-1.7B 56 of 56 (1,895), Bonsai-8B 56 of 56 (1,568); 12 of 12 refusals each. - The rest of the default chain (
penalty_oracle.py --kind sampler→oracle_samplers,oracle_samplers_8b): typical-p, top-n-σ, XTC, dynamic temperature and DRY, the same way, on 23 variants × 4 prompts — each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and the repeat penalty, all five at once. mindx-gen39 76 of 76 (3,576 tokens), Bonsai-1.7B 76 of 76 (2,587); 16 of 16 refusals each. The Bonsai-8B record is still to be taken (CHANGELOG 0.3.7). std::sort's order (testing/sort_oracle.cpp→oracle_std_sort). typical-p sorts withstd::sort, which is not stable, so the order equal scores end in is the algorithm's own. The oracle is libstdc++ itself, compiled by the gate: 876 of 876 orders identical, sizes 0 to 1,000, heavy ties, sorted, reversed and equal keys.- Live (
--bankml, the gate'spenalty_oracle_liveandsampler_oracle_live): every recorded case through/v1/chat/completionsand every fourth also through Ollama's/api/chat, each from an empty slot; a refusal must be a 400 with llama-server's message. The penalties: 85 of 85 on mindx-gen39. On/api/chatonly typical-p of the 0.3.7 samplers exists; the others are not Ollama options and go through/v1.
5g. The serving contract (0.3.8, unreleased)
Each records llama-server b11192's answers (--record) and replays the same requests through bankml serve --native
on Bonsai-1.7B.
- The context limit (
testing/context_oracle.py→context_oracle_live). Both servers at a 256-token context, with context shift off (llama-server's default). Eight requests from a few tokens to past the context: the same text, counts andfinish_reason("length"at the full context), and for a prompt that does not fit the same 400 status and body (exceed_context_size_error). 8 of 8. - Slots (
testing/slot_oracle.py→slot_oracle_live).POST /slots/0?action=save|restore|erase. Here the answer to compare with is the engine's own: an answer after a restore, in the same server and after a restart, must equal the answer computed from an empty slot, and the restore must skip the prompt (cache_n> 0). By the rule at the end of this file that half is a regression test. The refusals are llama-server's (501 without a slot directory, "Invalid slot ID", "Invalid action", "Invalid filename", "Unable to restore slot: …"), and a file from another model or damaged by one bit is refused. 19 of 19. - One slot and the host prompt cache (
testing/session_oracle.py→session_oracle_live). Three conversations take turns (A1 B1 C1 A2 …) through llama-server with its prompt cache on, so each turn finds another conversation's tokens in the slot: bankML must give the same text, counts andcache_n, turn by turn (modules/prompt_cache.md). Then four requests sent at once must all be answered, none mixed, each with the text it gets alone; llama-server's order among simultaneous arrivals is not deterministic, so this half checks bankML against itself. 14 of 14. - Logprobs (
testing/logprobs_oracle.py→logprobs_oracle_live)./v1withlogprobsandtop_logprobs: the same token ids, texts and bytes, every logprob the same 32-bit float, llama-server's entry rules for stop words and split UTF-8, and its refusals. Five cases are streamed, and there every chunk's delta and entry must match. 14 of 14.
5h. The q8_0 KV cache and the grammar trie (0.3.9, in progress)
- The KV cache's kernels (
oracle_ggml_b11192_q8_0_kv_kernels,forward.rs). The shipped haswell library, loaded withdlopen:quantize_row_q8_0byte for byte andggml_vec_dot_q8_0_q8_0bit for bit, on random activations with outliers and on blocks of −128 and 127 wheremaddubssaturates. 4,000 rows quantized byte-exact, 4,000 dot products bit-exact. - Answers over the cache (
testing/kv_oracle.py→kv_oracle_live). llama-server b11192 with--cache-type-k q8_0 --cache-type-v q8_0, and bankML withBANKML_CACHE_TYPE=q8_0: greedy and seeded, a 2,244-token prompt, a 320-token answer and a two-turn conversation; text, counts andcache_n. 6 of 6 (567 tokens). The first replay passed 2 of 6: llama.cpp rotates Q, K and V through a Hadamard transform around a quantized cache, which bankML now reproduces (fwht). - The grammar mask through a trie (
oracle_grammar_masks, §5c). The oracle now computes every mask both ways, one candidate at a time and by the trie's walk, each against llama.cpp's mask: 196 of 196 runs, 1,645 of 1,645 masks identical to llama.cpp b11192 by each.
6. Oracles planned
- P3, bankml's own forward pass: done. The criterion fixed here before it was built — at temperature 0, on the same prompts, an answer token-identical to llama.cpp b11192's — is met by §1d's greedy oracles (0.2.7 on), under seeded sampling too (0.2.11), and for whole conversations (0.3.0).
- Planned (TODO.md): tool calls, each with llama-server's answers; a 4-bit (
q4_0) KV cache with the same rotation;Q8_0weights with an oracle for Qwen's own template; NEON kernels against llama.cpp's ARM build. - Speculative decoding (measured in 0.1.8). A draft model (Bonsai-1.7B) and n-gram speculation both left the output token-identical at temperature 0 on every run; neither was faster beyond this laptop's noise, so neither is the default (n-gram is an opt-in). The same criterion applies to any future speed-up that changes how tokens are computed.
- The upstream
Q2_0kernel (0.2.0).upstream/test_q2_0_avx2.cholds the C drop-in to the shippedggml_vec_dot_q2_0_q8_0, bit for bit: 200,000 of 200,000 random cases.
The rule, restated
- An oracle is external: a compiled library, a standard's vectors, a published value, an independent implementation. A test that compares bankml with bankml is a regression test, not an oracle.
- A comparison is exact where the quantity is exact: bits, bytes, hashes. Tolerances are used for nothing that an oracle can check exactly.
- A speed claim is made only on code that passed every oracle in the same gate run, and only if the gain is beyond the run-to-run noise measured on the same machine.
- Every gate run is kept (
testing/results/<version>.txt), including the rejected experiments' numbers.