bankml / docs /oracles.md
Gregory-L's picture
bankML: the whole source (github.com/cryptoAGI/bankml @ 12ae409) and its page, with the bankML persona; the live engine (Dockerfile, hf/start.sh) ready for Docker hardware
28c70af verified
|
Raw History Blame Contribute Delete
50.3 kB

Oracles in bankml

What an oracle is here

An oracle is an answer bankml did not compute and cannot influence, against which its own answer is compared. bankml's rule since 0.0.1 is that a result counts only if an oracle has checked it: the same bits first, then the speed. A faster kernel that is not bit-identical to its oracle is not a result; a construction (a hash, a Merkle root, an ABI encoding) that does not reproduce a published value is not trusted.

The strongest oracle is the compiled reference itself: llama.cpp b11192's own shared libraries, called in-process through their exported symbols on the same bytes bankml reads. The others are published values: test vectors from standards (FIPS 180-4, RFC 6962), values published by the systems bankml interoperates with (Savante's doctrine root, the THOT spec's vectors, Hugging Face's and Ollama's sha256 for a model file), and independent implementations (pycryptodome, Foundry's cast, the Python guard bankml's Rust guard was ported from).

This file lists every oracle bankml uses, what it checks, how to run it, and what it last found. The results are in testing/results/, one record per release, written by testing/release_gate.sh. The latest record is 0.3.6's; figures marked 0.3.7, 0.3.8 or 0.3.9 are from CHANGELOG.md for versions not yet released, and reach testing/results/ when each is cut. Each module's own oracles are also listed on its page in modules/, under How it is verified.

The release gate, stage by stage

Every stage of testing/release_gate.sh, in the order it runs. A stage that fails stops the gate. The rows marked regression compare bankml with itself or with its own scalar model; they guard against change, and by the rule at the end of this file they are not oracles.

Always (no model needed)

stage checks against §
cargo test, cargo test -p bankml-capi unit tests, the scalar models of §2, the gateway's receipts against a mock server scalar models of ggml; a mock llama-server (regression where it is bankml against bankml) §2, §5
cargo clippy -D warnings lint, both packages — —
spdx_check.py every tracked source file names its licence in an SPDX header matching its layer LICENSING.md —
test_gguf_guard.py the Python guard's own suite its fixtures §3
test_ui.py the UI's data layer: keccak256, THOT manifests, CIDv1, the RFC 6962 tree published values §4
test_console.py (0.3.7) the console's routes, loopback and CSP rules, the SELF block's units, the persona's doctrine root, the Infotags metadata RFC 6962 roots; regression otherwise §4
test_connectors.py PostgreSQL publish and load on a throwaway cluster; a tampered row refused the manifest's bytes §4
test_chain.py iNFT mint and load on a throwaway anvil devnet the compiled contract §4
test_models.py the model importer: the pin, the licence gate, resume, a carrier started and rolled back published sha256 values §4
guard_agree.py the Rust guard against the Python guard, JSON for JSON the Python original §3
capi_oracle.py --printf bankml_log byte for byte libc snprintf §5b

With the b11192 release (BANKML_GGML_LIB), each a Rust test run by name (cargo test --release -- --ignored --exact), then the live checks against a running bankml serve --native:

stage checks against §
schema_oracle.py, content_oracle.py (when LLAMA_SRC is set) re-record the schema grammars and the content rule llama.cpp's own code in libllama-common.so §5e
sort_oracle.cpp (when g++ is present) re-record std::sort's orders libstdc++ §5f
oracle_tokenizer token ids, 4,346 cases llama-server /tokenize §1b
oracle_chat_template prompts byte for byte, 317 conversations llama-server /apply-template §1c
oracle_forward_embed_norm, oracle_forward_qkv_rope, oracle_forward_attention, oracle_forward_attention_tiled, oracle_forward_attention_split, oracle_forward_swiglu_sweep layer 0 operation by operation, the three attention kernels, ggml's expf the shipped ggml §1d
oracle_forward_model, oracle_forward_model_ternary every layer, result_norm and the logits the shipped ggml graph §1d
oracle_greedy_llama_server, _ternary, _long, _deep greedy tokens, short, long and deep prompts llama-server §1d
oracle_sample_llama_server seeded sampling, 40 continuations llama-server §1d
oracle_native_serve three conversations, 9 turns: text, counts, cache reuse llama-server /v1 §1d
oracle_grammar_masks whole-vocabulary masks, both paths since 0.3.9 libllama's grammar sampler §5c, §5h
oracle_json_mode, oracle_json_mode_ternary answers under JSON mode and user grammars llama-server §5c
oracle_schema_grammars, oracle_json_schema, oracle_json_schema_ternary, oracle_json_schema_o4, oracle_json_content schema grammars, answers under schemas, the content rule llama.cpp's own code; llama-server §5e
oracle_persona_layer mindX's persona Modelfile, two ways llama-server with the persona as system message §5e
oracle_penalties, oracle_penalties_8b the repeat, frequency and presence penalties llama-server §5f
oracle_samplers, oracle_samplers_8b (0.3.7) typical-p, top-n-σ, XTC, dynamic temperature, DRY llama-server §5f
oracle_std_sort (0.3.7) std::sort's order of equal keys libstdc++ §5f
oracle_ggml_b11192_q8_0_kv_kernels (0.3.9) the q8_0 KV cache's quantizer and dot product the shipped ggml §5h
gpu_q1_0_mat_vec_bit_exact the GPU kernels on every usable card the CPU kernel, itself proven against ggml §1d (0.2.13)
oracle_train_script, oracle_train_imprint mindXtrain's author and score stages mindXtrain's Python §1d (0.2.13)
oracle_ggml_b11192_real_bonsai_1_7b, _real_bonsai_8b_q1_0, _real_ternary_bonsai_8b every weight, q8_0 rows, dot products the shipped ggml §1
oracle_ggml_b11192_f16 F16 products, both of ggml's paths the shipped ggml §5d
oracle_tokenizer_smollm, oracle_chat_template_chatml, oracle_forward_model_bonsai_1_7b, oracle_forward_model_llama_f16, oracle_llama_server_bonsai_1_7b, oracle_llama_server_llama_f16, oracle_native_serve_o4 the O4 models: tokenizer, templates, whole model, llama-server's tokens, conversations and JSON mode llama-server; the shipped ggml §5d
ab_vs_ggml, ab_vs_ggml_q2_0, bench_q1_0_prefill_act, bench_memory_floor, decode_budget_q1_0, decode_budget_q2_0 speed: kernel A/Bs (each first checks agreement), the memory floor, one token's matmul budget the shipped ggml's own kernels; measurements, not oracles §1, PERFORMANCE.md
serve_oracle_ollama_shape the conversations through /api/chat, /v1, and /v1 after a reload llama-server's record §1d (0.3.1)
o4_live the conversations, JSON mode and schemas live, on Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 llama-server's records §5d, §5e
persona_oracle_live the created mindx-gen39 through /api/chat, /api/generate, /v1; /api/ps llama-server's record §5e
penalty_oracle_live, sampler_oracle_live the penalties and the sampler chain through /v1 and /api/chat llama-server's records §5f
context_oracle_live (0.3.8) the context limit llama-server's record §5g
slot_oracle_live (0.3.8) slot save, restore and erase the engine's own empty-slot answer; llama-server's refusals §5g
session_oracle_live (0.3.8) interleaved conversations; simultaneous requests llama-server's record; the queue against itself §5g
kv_oracle_live (0.3.9) answers over a q8_0 KV cache llama-server with --cache-type-k/v q8_0 §5h
logprobs_oracle_live (0.3.8) logprobs, streamed and not llama-server's record §5g
json_oracle_live, json_schema_oracle_live JSON mode and schemas live on the 8B model llama-server's records §5c, §5e
capi_chat_oracle bankml_chat serve --native; llama-server's record §5b

When their inputs are present

stage checks against §
oracle_convert_b11192 (per directory under .models/convert/) bankml convert's GGUF, sha256 for sha256 b11192's convert_hf_to_gguf.py §5e
oracle_name_heuristics (when BANKML_LLAMA_SRC is set) model-name heuristics, 168 ids gguf-py §5e

1. The ggml oracle: kernels bit-exact against the compiled llama.cpp

What it checks

For each real model, three things, every one to the bit:

# quantity bankml side ggml side (llama.cpp b11192, shipped binaries)
1 every weight of every low-bit tensor, dequantized to f32 q1_0::dequantize_row, q2_0::dequantize_row dequantize_row_q1_0 / dequantize_row_q2_0 in libggml-base.so
2 activation rows quantized to q8_0 (32 signed bytes and a half-precision scale per block) quantize_row_q8_0 quantize_row_q8_0 in libggml-cpu-haswell.so
3 dot products of a weight row with a q8_0 row vec_dot_ref (scalar model), vec_dot / vec_dot_act (AVX2), vec_dot_act_scalar ggml_vec_dot_q1_0_q8_0 (AVX2 path), ggml_vec_dot_q1_0_q8_0_generic, ggml_vec_dot_q2_0_q8_0

Dequantized tensors are compared by the sha256 of their little-endian f32 output (a whole 8-billion-weight model, tensor by tensor, without holding it in memory); quantized rows byte for byte; dot products by their f32 bit patterns (to_bits()), not within a tolerance.

Why "the compiled library" and not "the source"

The first port of ggml's generic C, done faithfully from the source, disagreed with the shipped library in the last bit. Reading the disassembly showed why: the compiler had fused a multiply and an add into one FMA instruction, which rounds once instead of twice. A source-level port cannot see that; the binary oracle catches it. bankml's kernels reproduce the float order and the fused multiply-adds of the shipped libggml-cpu-haswell.so (the variant ggml's backend scorer loads on an AVX2 CPU without AVX-512), and the oracle proves it on real tensors.

It also shows the oracle can tell float orders apart. For the ternary format ggml ships two builds of the same generic C: the haswell variant (FMA) and the baseline x64 variant (SSE, no FMA). They disagree with each other on 19 of the 762 recorded ternary dot products. bankml carries a model of each (vec_dot_ref for haswell, vec_dot_ref_nofma for x64) and matches each one in 762 of 762.

How it runs

  1. Record (once per model): testing/ggml_oracle.py loads the release's libggml-base.so and libggml-cpu-haswell.so (and, for Q2_0, libggml-cpu-x64.so) with ctypes, calls ggml_cpu_init(), memory-maps the GGUF and writes:

    • dequant.tsv: tensor name · element count · sha256 of ggml's f32 output, for every tensor of the type;
    • vecdot.bin: for sampled rows of every tensor (3 per tensor by default): the f32 activation x, ggml's q8_0 of it, and ggml's dot products (AVX2 and generic).

    The release tarball is checked by sha256 before use (llama-b11192-bin-ubuntu-x64.tar.gz, 34cf6fa5…81ec7).

  2. Compare (every gate): the Rust tests oracle_ggml_b11192_real_* (in q1_0.rs and q2_0.rs, #[ignore]d because they need the models) re-derive every recorded quantity from the same file through bankml's own memory map and assert equality.

  3. A/B against the live library (every gate): ab_vs_ggml and ab_vs_ggml_q2_0 dlopen the shipped haswell library directly (no crate, dlopen/dlsym declared by hand) and time ggml's kernel and bankml's on the same real weights in one process, after first checking they agree.

What it found (0.1.7–0.2.0 gates; identical in each, and again under Rust 1.99 in 0.3.2)

The gate runs these as oracle_ggml_b11192_real_bonsai_1_7b, oracle_ggml_b11192_real_bonsai_8b_q1_0 and oracle_ggml_b11192_real_ternary_bonsai_8b.

model tensors weights q8_0 rows dot products
Bonsai-1.7B, Q1_0 197 1,719,904,256 788 byte-exact 788 bit-exact (AVX2 and vec_dot_act); generic port == ggml generic 788/788
Bonsai-8B, Q1_0 254 8,188,239,872 762 byte-exact 762 bit-exact; generic 762/762
Ternary-Bonsai-8B, Q2_0_g64 254 8,188,239,872 762 byte-exact 762 bit-exact vs haswell (reference, AVX2 and scalar paths); no-FMA model == x64 build 762/762

Run it

# once: the oracle files (numpy + the b11192 release)
python3 testing/ggml_oracle.py .models/Bonsai-1.7B-Q1_0.gguf /path/to/llama-b11192 .models/oracle
python3 testing/ggml_oracle.py .models/Bonsai-8B-Q1_0.gguf /path/to/llama-b11192 .models/oracle-8b-q1
python3 testing/ggml_oracle.py .models/Ternary-Bonsai-8B-Q2_0_g64.gguf /path/to/llama-b11192 .models/oracle-ternary
# every time
cargo test --release -- --ignored oracle_ggml_b11192 --nocapture --test-threads=1
BANKML_GGML_LIB=/path/to/llama-b11192 cargo test --release -- --ignored ab_vs_ggml --nocapture --test-threads=1

1b. The tokenizer oracle: token-identical to llama.cpp (0.2.1)

P3, bankml's own forward pass, starts with the tokenizer. testing/tokenizer_oracle.py asks a running llama-server (b11192, the Bonsai / Qwen3 vocabulary) to tokenize a corpus, with special tokens parsed and not, and records every answer. The corpus is:

  • every document in the repository and Savante's canon texts;
  • the chat template's markers;
  • hand-picked edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with joiners, mathematical alphanumerics;
  • a seeded fuzz set of 2,000 strings drawn from 17 Unicode ranges.

tokenizer::tests::oracle_tokenizer re-derives every case with bankml's tokenizer (tokenizer.rs, no crates) and requires the same ids in the same order. Last result: 4,346 of 4,346 cases token-identical (the corpus grew at 0.3.4; 4,258 at 0.2.1), on the Qwen2 and the SmolLM2 pre-tokenizer each.

Two details the oracle settles:

  • Which special tokens split the text. Qwen3's <think> markers are USER_DEFINED and split the text even when special tokens are not parsed; CONTROL tokens such as <|im_start|> do not.
  • The pre-tokenizer's letter class. It is Unicode general category L, which is not char::is_alphabetic. It is generated into unicode_letters.rs.

1c. The chat-template oracle: byte-identical prompts (0.2.2)

A conversation becomes a prompt through the model's chat template, a Jinja program inside the GGUF. bankml does not run Jinja. chat.rs writes the Bonsai / Qwen3 template's rules out, and check_template accepts only that template, identified by the sha256 of its text. testing/template_oracle.py records llama-server's own /apply-template for 317 conversations. They cover:

  • a system prompt first, later, or absent;
  • assistant turns with and without <think> blocks, before and after the last real user query;
  • reasoning_content;
  • runs of tool results;
  • user messages that look like tool responses;
  • special markers and Unicode inside content;
  • 300 random conversations.

oracle_chat_template requires every prompt byte-identical: 317 of 317. The oracle found one server behaviour that the template alone would not predict: an empty reasoning_content is dropped before templating.

1d. The forward-pass oracle: llama.cpp's own operations, from the shipped ggml (0.2.3)

The release has no tool that prints a model's intermediate values, so the oracle builds them the way the kernel oracle does. testing/forward_oracle.py drives the shipped libggml through ctypes with the same operations llama.cpp's Qwen3 graph uses and computes them with the release's own CPU backend. For 300 real token ids (the chat markers, the table's edges, and tokens of the tokenizer's corpus) it records the sha256 of each row's f32 bytes:

  • get_rows on the Q1_0 token embedding: llama.cpp's inp_embd;
  • rms_norm with the model's epsilon, then mul by blk.0.attn_norm.weight: llama.cpp's attn_norm-0.

bankml's forward.rs computes the same in the float order read from b11192's ops.cpp:

  • the sum of squares accumulated in double, one float product at a time, in order;
  • the mean rounded to float;
  • scale = 1 / sqrtf(mean + eps);
  • each output (x · scale) · w, with no FMA.

oracle_forward_embed_norm requires 300 of 300 rows bit-exact for both.

Step four (0.2.4): Q, K, V, the head norms and RoPE. The oracle also builds layer 0's attention inputs on a 28-token chat prompt:

  • mul_mat of the Q1_0 attn_q, attn_k and attn_v weights with the normed rows;
  • reshape to 128-wide heads, rms_norm and mul by attn_q_norm and attn_k_norm;
  • rope_ext in NEOX mode with the YaRN parameters llama.cpp's context derives (freq_scale 0.25, ext_factor 1, attn_factor 1.0, whose bits the record carries, beta 32 and 1, n_ctx_orig 16,384, base 10⁶).

RoPE runs twice: at positions 0–27, and at positions 7 to 63,214, where theta is a long product. oracle_forward_qkv_rope requires 140 of 140 rows bit-exact.

This step showed why the oracle is the shipped binary and not the source. Written as ops.cpp reads, bankml's RoPE matched only 30 of 84 rows, while the projections and norms before it were exact. The disassembly of libggml-cpu-haswell.so shows GCC's FMA contraction of the rotation, the YaRN mix and the magnitude term. With those three written as the same mul_adds, every row matches.

Step five (0.2.5): attention. The oracle stores K and V through cpy to f16, as llama.cpp's cache does. It builds the causal f16 mask and runs flash_attn_ext, which is llama-server's default attention on the CPU, with F32 accumulation set as llama-graph sets it. The output then goes through wo (mul_mat) and the residual add. oracle_forward_attention requires 84 of 84 rows bit-exact (kqv_out, attn_out, ffn_inp, 28 tokens).

The reproduction follows ggml's reference path:

  • the f16 dot's four 8-lane FMA accumulators and their reduction tree;
  • an online softmax whose V accumulator is rounded to f16 at every step.

The oracle rejects a near miss: with the softmax sum written as an FMA, only 59 of 84 rows match. At 0.2.5 two other kernels were not yet covered; each got its own oracle in 0.2.9 and 0.2.10 (below):

  • the tiled kernel, for 64 or more query rows;
  • the split-KV kernel, for a decode whose padded KV length reaches 512 (llama.cpp pads to multiples of 256, so from 257 cells in use).

Step six (0.2.6): the feed-forward block; layer 0 complete. The oracle continues from ffn_inp:

  • rms_norm and mul by ffn_norm;
  • mul_mat by ffn_gate and ffn_up;
  • swiglu_split;
  • mul_mat by ffn_down;
  • add, giving l_out, the layer's output.

oracle_forward_attention requires 112 of 112 rows through l_out. SiLU in the shipped build uses ggml's own vectorized expf, a polynomial in FMAs, not libm. bankml reproduces it lane by lane.

A second check, oracle_forward_swiglu_sweep, feeds 24,600 values across ±120 and the edges straight to ggml_swiglu_split and requires every output's bits (24,600 of 24,600). The sweep exists because the layer check alone was not enough: with libm's expf, l_out still matched 27 of 28 rows, since q8_0 quantization absorbs most of the difference, while the sweep matched only 19,245 of 24,600.

Step seven (0.2.7): the whole model, and llama-server itself.

  • testing/model_oracle.py computes the whole Qwen3 graph in the shipped ggml: every layer, output_norm and mul_mat by output.weight. It runs one layer per ggml context and carries the residual stream between contexts as f32 bytes, which is the same arithmetic as one graph in a fraction of the memory. oracle_forward_model requires 1,064 of 1,064 rows bit-exact: 36 × 28 l_out rows, 28 result_norm rows, and 28 logit rows of 151,669 values each.
  • testing/greedy_oracle.py steps outside the graph. It asks the running llama-server to render 6 chat prompts (/apply-template), tokenize them (/tokenize) and continue them greedily (/completion, top-k 1, prompt cache off). oracle_greedy_llama_server requires bankml's own forward pass to produce the same tokens on every prompt (6 of 6, 164 tokens), and to end the turn where the server did.

Both stayed inside the range bankml reproduced at 0.2.7: prompts under 64 tokens, contexts up to 256 cells (the largest case uses 80). Steps nine and ten removed both limits. llama.cpp pads the KV length to multiples of 256, so a decode beyond 256 cells takes the split-KV kernel.

Step eight (0.2.8): the ternary model. Both whole-model checks run again on Ternary-Bonsai-8B (Q2_0_g64), whose every matrix, including the embedding and the output, is ternary:

  • oracle_forward_model_ternary checks bankml against the shipped ggml graph: 1,064 of 1,064 rows.
  • oracle_greedy_llama_server_ternary checks against llama-server running the ternary model, with Savante's flags, on a spare port: 6 of 6 prompts, 140 tokens.

The greedy comparison now covers exactly the tokens the server generated. Its list includes the end-of-turn token when it produced one. Where it stopped otherwise, that is its stopping policy, not a token choice.

Step nine (0.2.9): long prompts and ggml's tiled kernel. llama.cpp computes a prompt in micro-batches of up to 512 tokens. A micro-batch of 64 rows or more takes flash_attn_ext_tiled, a different algorithm from the reference path: f32 Q, a SIMD GEMM per 64-cell KV tile, a vectorized softmax summed in double, and an f32 accumulator.

  • The forward oracle runs layer 0 on a 150-row micro-batch in the shipped ggml. oracle_forward_attention_tiled requires 150 of 150 rows bit-exact. The reference kernel would match only 1, the first row, so the kernel choice is itself checked.
  • greedy_oracle.py --long records llama-server's continuations for 6 prompts of 111–116 tokens. oracle_greedy_llama_server_long requires the same tokens on all 6.

Step ten (0.2.10): long contexts and ggml's split-KV kernel. A single-token decode over 512 or more padded cells cuts the cells into one chunk per llama.cpp thread and reduces the partial softmaxes, so the thread count shapes the bits.

  • The forward oracle records one decode row at 7 context lengths (257–1,000 cells) and at 3 and 4 threads. oracle_forward_attention_split requires 14 of 14.
  • greedy_oracle.py --deep records 3 continuations of 200 tokens that run to about 300 cells. oracle_greedy_llama_server_deep requires the same 600 tokens.

The 4-thread cases did double duty. Besides showing the dependence on the thread count, they exposed the reduction's FMA contraction: at 3 threads the first chunk always held the maximum, where the fused and plain forms agree.

Step eleven (0.2.11): sampling. testing/sample_oracle.py has llama-server sample 40 continuations with fixed seeds, and keeps the parameters the server reports it ran (generation_settings). oracle_sample_llama_server replays them through bankml's forward pass and sampler.rs and requires every token (40 of 40, 1,175 tokens). The cases range over temperature 0–1.5, top-k 5–128, top-p and min-p. Top-k is libstdc++'s partial_sort, ported line for line, because the order it leaves tied logits in decides which token a draw lands on.

0.2.13: the GPU and mindXtrain.

  • gpu_q1_0_mat_vec_bit_exact runs both Q1_0 GPU kernels on every usable card against the CPU kernel, which is bit-exact against ggml: on the Vega 3, every row of every shape from 64×128 to 12288×4096. bankml gpu --verify runs the same check for a user, and bankml trusts a card only after it passes.
  • oracle_train_script checks bankml's author stage against mindXtrain's scripts.py on 84 scripts, byte for byte.
  • oracle_train_imprint checks the score stage against score_imprint on 3,000 randomized cases, value for value.

0.2.14: the GPU inside the forward pass. With a verified card present, Weights::open gives it a share of every 1-bit matrix's rows, so every token oracle runs with the GPU working. The whole-model check passes at 1,064 of 1,064 rows with the Vega 3 on 26 % of every matrix.

The on-card oracle learned the lesson of this release. It now includes layer-shaped data: scales and magnitudes that vary per block, so products are inexact in f32. On that data the card's unfused Fma differed by one ulp, where random data had let it pass. bankml now computes the FMA exactly (Boldo–Melquiond), and with the driver's Fma put back, bankml gpu --verify refuses the card.

0.3.0: whole conversations through the server. testing/serve_oracle.py holds three Savante-style conversations, three turns each, with fresh seeds at temperature 0.3. It sends them to a fresh llama-server's /v1/chat/completions, the path Savante uses, so the server's prompt cache carries each conversation from turn to turn, and records every answer with its prompt count, completion count and the prompt tokens taken from the cache. oracle_native_serve replays the conversations through bankML's native engine in one process and requires all of it on 9 of 9 turns: the same text, the same counts, the same cache reuse. The rule it reproduces is visible in the numbers. Turn 2 reuses 49 of turn 1's 53 prompt tokens, because the empty <think> block that ended turn 1's prompt is not rendered when that answer becomes history. A new conversation reuses the 38 tokens of the shared system prompt.

0.3.1: the Ollama shape. testing/serve_oracle.py --bankml starts bankml serve --native with the 8B 1-bit model on spare ports and sends the same conversations three ways, each from an empty slot (the model unloaded, then loaded again, so each run starts as a fresh llama-server does): through Ollama's /api/chat, with the first conversation streamed as NDJSON and the options mapped from Ollama's names; through /v1/chat/completions; and through /v1 once more after an unload in which /v1 itself loads, and verifies, the startup model. Every turn must give the same text, prompt and completion counts and cache reuse on every path, equal to llama-server b11192's record. So the Ollama path inherits the conversation oracle, and a reload changes nothing. The gate runs it as serve_oracle_ollama_shape.

Every step of the forward pass above was added to these oracles before it counted, and every later feature (§5b–§5h) follows the same rule.

2. Scalar models: every fast path against its own reference

Between the real-model oracle runs, the kernels are held to a scalar model of ggml, on synthetic inputs, in the ordinary cargo test:

check where cases
half-precision conversion, all 65,536 values, and against the CPU's F16C instruction q1_0.rs 65,536
q8_0 rounding (round-half-to-even, scale = 127 / max|x|) q1_0.rs edge cases
AVX2 Q1_0 dot == scalar model of ggml q1_0.rs 3,500
AVX2 and portable Q2_0 dot == ggml model q2_0.rs 3,300
matrix–vector, matrix–matrix and threaded paths == per-pair dots q1_0.rs, q2_0.rs, par.rs every shape tested
the 0.0.4 prefill tile == per-pair q1_0.rs (act_tile_bit_exact) every tile
the generic port == f64 arithmetic on dequantized data (the mathematical definition) q1_0.rs random blocks

A speed-up is admitted only after these and §1 pass on the same build. Kernels that were faster but not exact, or exact but not faster beyond noise, are kept with their numbers in testing/experiments/ (0.0.5).

3. The guard: Rust against the Python it was ported from

bankml's header guard (gguf.rs: play, refuse or need more bytes; the three low-bit traps; hostile headers) was ported from minaiml's Python gguf_guard.py, vendored in testing/gguf_guard.py with its suite. testing/guard_agree.py runs both on every synthetic case and on every real model present and requires identical JSON. Last result: 28/28 agree.

4. Cryptographic constructions against published values

construction oracle where
SHA-256 (the model pin) FIPS 180-4 test vectors; since 0.1.8 the SHA-NI path must also equal the portable rounds on every length 0–1,000 (and 4 KiB, 64 KiB, split updates), and a real 1.16 GB model must hash to coreutils sha256sum's value and its published pin sha256.rs (fips_vectors, hardware_path_equals_portable_on_every_length)
the model pin the sha256 the publisher lists: the FORK.json of the PYTHAI fork, a Hugging Face repository's LFS sha256 at a fixed revision, or the Ollama registry's layer digest; every import is hashed as it streams and kept only if equal bankml.rs (pin), sAGI/models.py
keccak256 (pure Python) pycryptodome's keccak on every input length 0–400 bytes; and Savante's published doctrine root 0x92fe83eb…ae137d0, reproduced from her persona sAGI/agents.py, testing/test_ui.py
THOT manifests (sagi.thot_manifest/1) the spec's own test vectors (THOT_MANIFEST.md §5: savante@1fcca89, jaimla@8b57ccf, luvai@0c1eef7): bundle root, Merkle root, identity CID sAGI/thot.py, testing/test_ui.py
CIDv1 (raw, sha2-256, base32) the CID of "abc" that mindX's rage.py and Savante's ledger construction give testing/test_ui.py
.history Merkle tree (RFC 6962) the Certificate Transparency reference roots for 1, 3 and 8 leaves testing/test_ui.py
iNFT ABI encoding the compiled iNFT_7857 contract itself, deployed from its artifact on a throwaway anvil devnet: it must accept the calldata bankml encodes (simulate and mint), refuse what it should (missing role, a repeated content root), and read back exactly the values encoded sAGI/chain.py, testing/test_chain.py
Savante's canon her own offline verifier bind/savante_verify.py: 12 of 12 commitments, doctrine root the Verifier tab

5. The gateway: receipts against the text

bankml serve's receipts are checked end to end against a mock llama-server whose answers are known (testing/cli.rs): the response_sha256 equals the sha256 of the text the mock sent (streamed and not), request_sha256 equals the sha256 of the request body, and a model file changed after verification produces a 503 and no receipt. In use, the UI recomputes every answer's sha256 and marks ✓ or ≠ received!.

5b. The C API (0.3.2): the library against libc, and against the server

Two oracles hold the C API (capi/, CAPI.md), both run from C programs compiled with the system cc.

  • bankml_log against libc snprintf (testing/capi/printf_oracle.c). bankml_log is a C-variadic function defined in Rust, and its formatter is bankML's own, so the oracle is the C library's snprintf. Each case passes the same format and arguments to both. The message delivered to the sink must equal snprintf's output byte for byte, length included. The cases cover:

    • integers at INT_MIN/LLONG_MIN/SIZE_MAX, in every length;
    • %f ties (%.0f of 0.5, 1.5 and 2.5), %.1100f of the smallest subnormal, DBL_MAX, and inf/nan with signs;
    • strings, NULL, and an unterminated buffer bounded by a precision;
    • %c of 0 and 255, %p, %%, widths, precisions and *.

    Then the cases it does not support must give their exact markers. No argument may be read that cannot be typed, and %n must not write. The same comparison runs in cargo test -p bankml-capi, from Rust.

  • bankml_chat against serve --native and llama-server (testing/capi/capi_oracle.py --chat, the gate's capi_chat_oracle). The conversation oracle of §1d, carried across the C boundary:

    • on Ternary-Bonsai-8B, the same request bytes go to a live serve --native and then to bankml_chat in a fresh process. They must give the same text, the same streamed pieces, the same counts and cache reuse, and the same receipt hashes;
    • on Bonsai-8B Q1_0, bankml_chat must give llama-server b11192's recorded turns;
    • Bonsai-1.7B and an unpinned file must be refused, with serve's reasons.

5c. JSON mode and grammars (0.3.3): llama.cpp's grammar sampler, and llama-server's answers

Two oracles hold O6's first cut (bankML/grammar.rs). Neither compares bankML with itself.

  • The grammar engine against llama.cpp's own (testing/grammar_oracle.py → oracle_grammar_masks). testing/grammar_oracle.cpp drives libllama b11192's public API (llama_sampler_init_grammar, _apply, _accept) on the pinned vocabulary, loaded vocab-only. The grammar functions inside llama.cpp are not exported, but the sampler that wraps them is, so the oracle is the code the server runs, not a model of it.
    • The vocabulary as the grammar reads it. Every token's piece (llama_token_to_piece, special tokens rendered) and whether it ends generation. bankML must agree on 151,669 of 151,669 tokens and on the end set of 6. That set includes <|fim_pad|>, <|repo_name|>, <|file_sep|> and token 128247 </s>, which bankML's engine had not counted as ends before.

    • Masks. The whole-vocabulary mask, before every token and after the last, is recorded as how many tokens are allowed and an FNV-1a hash of which. The runs span 12 grammars:

      • llama-server's JSON-mode grammar, with its generation-prompt prefill;
      • all 8 grammars in llama.cpp's grammars/ (json, json_arr, arithmetic, c, chess, english, japanese, list);
      • three written for what those miss: token terminals <think>, !<|im_end|> and <[151644]>; .; {m,n} edges; comments and CRLF; astral and CJK ranges.

      The inputs are valid, invalid, partial, escaped and unicode JSON, whitespace past space's 20-character and 2-newline limits, prose, fences and truncations. Each goes in twice: tokenized as the tokenizer does, and as one token per byte, which splits every multi-byte character across tokens and exercises the partial-UTF-8 path. The result: 196 of 196 runs; 1,645 of 1,645 masks and 116 of 116 rejection points identical.

  • llama-server's answers under a grammar (testing/json_oracle.py --record → oracle_json_mode, oracle_json_mode_ternary). llama-server b11192 runs with Savante's flags and --verbose, so that each answer reports its tokens and the grammar it ran.
    • What is sent. /v1/chat/completions with response_format: {"type": "json_object"} on 7 prompts, greedy and seeded at temperatures 0.7, 1.0 and 1.3 (top-k 40). Some prompts invite JSON. Others tempt prose, a haiku or code, so that the grammar has to reject what the model wants: 152 of 860 tokens on the 1-bit model were redrawn under the mask. One is a nested object of 159 tokens, one is unicode, and two answers are cut by max_tokens. The 1-bit model also gets two user grammar requests.
    • What is required. Each case runs from an empty cache. bankML's engine must give the same token ids, the end token included (one answer ends on <|file_sep|>), and the same raw text, content, finish reason and counts. The grammar text and generation prompt the server reports must be bankML's constant and prefill.
    • The result: Bonsai-8B Q1_0 23 of 23 (860 tokens); Ternary-Bonsai-8B 13 of 13 (534 tokens, 49 redrawn).
    • The live check (testing/json_oracle.py --bankml, the gate's json_oracle_live) sends a subset through a running serve --native, each from an empty slot: /v1 with response_format, streamed once; /v1 with grammar; and Ollama's /api/chat with format: "json". Every answer must equal the record: 8 of 8.

5d. O4 (0.3.4): F16 products, the Llama graph, tied embeddings — the same oracles, three new models

Every oracle family the 8B models have, extended to Bonsai-1.7B (tied embeddings), SmolLM2-135M-Instruct and mindx-gen39 (the Llama graph in F16), each against llama-server b11192 running that very file; plus one new kernel oracle.

  • F16 products against the shipped library (testing/f16_oracle.py → oracle_ggml_b11192_f16). ggml multiplies an F16 weight two ways: ggml_vec_dot_f16 for one column, llamafile's tinyBLAS for two or more (when the weight has a multiple of 4 rows and the row a multiple of 8). The oracle builds ggml_mul_mat graphs through ctypes and has the shipped libggml-cpu-haswell.so compute them: SmolLM2's real matrices (layers 0 and 29, and the 49,152-row token table) by 1, 2, 3, 5 and 28 columns, and synthetic shapes that reach the tails (k % 32 ≠ 0) and the fallbacks (rows % 4 ≠ 0, k % 8 ≠ 0). Every F16 tensor widened by ggml_fp16_to_fp32_row is hashed too. The result: 211 of 211 tensors; 552,268 of 552,268 elements in 87 products (61 tinyBLAS, 26 vec_dot). The two paths give different bits for the same row and column (the_two_reductions_differ), so the oracle checks the choice as well.
  • The whole model (testing/model_oracle.py, now for tied embeddings, F16 and the Llama graph: rope_ext mode 0, no Q/K norms, ext_factor 0): every layer's l_out, result_norm and the logits. An F16 model is replayed as one micro-batch with every row output, because that is what the oracle's graph computes. Bonsai-1.7B 840 of 840 rows, SmolLM2-135M-Instruct and mindx-gen39 800 of 800 each (oracle_forward_model_bonsai_1_7b, oracle_forward_model_llama_f16).
  • The tokenizer (tokenizer_oracle.py 127.0.0.1:PORT .models/oracle-tokenizer-smollm → oracle_tokenizer_smollm): the same corpus on SmolLM2's vocabulary, 4,346 of 4,346, with two new edge strings: bytes SmolLM2 has no token for, and plane-4 characters. Re-recorded on the Qwen3 vocabulary, the grown corpus caught a bug: </s> (NORMAL in the file, CONTROL in llama.cpp by its name) was taken as text in 2 of 4,346 cases. Fixed; 4,346 of 4,346.
  • The templates (template_oracle.py, written per model as cases-<stem>.jsonl → oracle_chat_template_chatml): SmolLM2-Instruct's and mindx-gen39's ChatML, 317 of 317 each.
  • llama-server's tokens (greedy short, --long, --deep; seeded; conversations; JSON mode, each recorded from that model's server; oracle_llama_server_bonsai_1_7b, oracle_llama_server_llama_f16, oracle_native_serve_o4): SmolLM2 6, 6, 3 and 40 of 40; mindx-gen39 6, 6, 3 and 40 of 40; Bonsai-1.7B 6 and 40 of 40; conversations 9 of 9 and JSON mode 23 of 23 on each. On a ChatML template llama-server builds a different JSON grammar (no <think> in its root); the record shows it, and bankML carries it per template.
  • Live and through the C API: serve_oracle.py --bankml STEM [NAME] and json_oracle.py --bankml STEM [NAME] on each model (mindx-gen39 asked for as mindx-gen39), and capi_oracle.py --chat against each model's record. The gate runs the live pair, with json_schema_oracle.py --bankml since 0.3.5, as its o4_live stage.

5e. JSON schemas and bankml create (0.3.5): llama.cpp's converter, parser and converter script, and the persona

  • The schema grammar (testing/schema_oracle.{cpp,py} → oracle_schema_grammars): llama.cpp b11192's own json_schema_to_grammar and common_chat_templates_apply, called inside the release's libllama-common.so (no model), on 173 schemas — llama.cpp's 81 test cases, Pydantic-shaped schemas like mindX's, edge cases and every refusal path — with the template read from each model that carries it (Bonsai's Qwen3, SmolLM2-Instruct's ChatML, mindx-gen39's ChatML). 148 grammars byte-identical bare and on the chat path, on each template; every refusal with llama.cpp's message (24, 20). The ChatML wrapping (no reasoning rules, another root) was read from this output.
  • The answers (testing/json_schema_oracle.py --record STEM → oracle_json_schema, oracle_json_schema_ternary, oracle_json_schema_o4): llama-server b11192 asked under schemas through all three request shapes, greedy and seeded, each from an empty cache, with answers cut by max_tokens inside an object, a top-level string and a number. Required: the same token ids (end token included), raw text, content, finish and counts, and the server's grammar and generation prompt are bankML's. Bonsai-8B 28 of 28, Ternary-Bonsai-8B 11 of 11, Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 56 of 56 each. Live (--bankml STEM [NAME], the gate's json_schema_oracle_live and o4_live): /v1 (one streamed) and /api/chat format: <schema>.
  • The content rule (testing/content_oracle.{cpp,py} → oracle_json_content): llama.cpp's common_chat_parse itself, with the parser llama-server builds for each request, on every prefix of every recorded constrained answer and on edge cases: 30,063 of 30,063 texts on each template (36 texts with a reasoning block, which no grammar admits after its prefill, are refused by llama.cpp's parser and listed). It found two differences the answer oracles had not: unfinished escapes, and llama-server's raw-text answer for an empty parse; both fixed.
  • The converter (testing/convert_oracle.py, oracle_convert_b11192, oracle_name_heuristics): bankml convert against llama.cpp b11192's convert_hf_to_gguf.py --outtype f16 on the same directory, sha256 for sha256 — gen39 6b64c748… and SmolLM2 ec30a679…, run directly in 0.3.5 (torch 2.14.1+cpu) — and gguf-py's name heuristics on 168 of 168 ids.
  • mindX's persona, end to end (testing/persona_oracle.py → oracle_persona_layer, and live as persona_oracle_live): promote.py's persona Modelfile created two ways (FROM the merged directory; FROM mindx-gen39 in place), the user's turns alone, against llama-server given the persona as the system message on the same GGUF: 27 of 27 token-identical for each; live through /api/chat, /api/generate and /v1 from an empty slot each, and /api/ps names the derived model. (Without the empty slot, seeded answers moved: a cached prefix changes F16's product paths, so the live check starts each answer as the record did.)

5f. The penalties (0.3.6) and the rest of the sampler chain (0.3.7, unreleased)

  • The penalties (testing/penalty_oracle.py → oracle_penalties, oracle_penalties_8b). llama-server b11192 runs with Savante's flags from an empty cache, greedy and seeded, on 17 variants × 4 prompts made to repeat; the window reaches into the prompt, as llama-server fills it. Each case keeps the parameters the server read back. bankML must give the same tokens, and refuse what the server refuses with its message. mindx-gen39 56 of 56 (2,478 tokens), Bonsai-1.7B 56 of 56 (1,895), Bonsai-8B 56 of 56 (1,568); 12 of 12 refusals each.
  • The rest of the default chain (penalty_oracle.py --kind sampler → oracle_samplers, oracle_samplers_8b): typical-p, top-n-σ, XTC, dynamic temperature and DRY, the same way, on 23 variants × 4 prompts — each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and the repeat penalty, all five at once. mindx-gen39 76 of 76 (3,576 tokens), Bonsai-1.7B 76 of 76 (2,587); 16 of 16 refusals each. The Bonsai-8B record is still to be taken (CHANGELOG 0.3.7).
  • std::sort's order (testing/sort_oracle.cpp → oracle_std_sort). typical-p sorts with std::sort, which is not stable, so the order equal scores end in is the algorithm's own. The oracle is libstdc++ itself, compiled by the gate: 876 of 876 orders identical, sizes 0 to 1,000, heavy ties, sorted, reversed and equal keys.
  • Live (--bankml, the gate's penalty_oracle_live and sampler_oracle_live): every recorded case through /v1/chat/completions and every fourth also through Ollama's /api/chat, each from an empty slot; a refusal must be a 400 with llama-server's message. The penalties: 85 of 85 on mindx-gen39. On /api/chat only typical-p of the 0.3.7 samplers exists; the others are not Ollama options and go through /v1.

5g. The serving contract (0.3.8, unreleased)

Each records llama-server b11192's answers (--record) and replays the same requests through bankml serve --native on Bonsai-1.7B.

  • The context limit (testing/context_oracle.py → context_oracle_live). Both servers at a 256-token context, with context shift off (llama-server's default). Eight requests from a few tokens to past the context: the same text, counts and finish_reason ("length" at the full context), and for a prompt that does not fit the same 400 status and body (exceed_context_size_error). 8 of 8.
  • Slots (testing/slot_oracle.py → slot_oracle_live). POST /slots/0?action=save|restore|erase. Here the answer to compare with is the engine's own: an answer after a restore, in the same server and after a restart, must equal the answer computed from an empty slot, and the restore must skip the prompt (cache_n > 0). By the rule at the end of this file that half is a regression test. The refusals are llama-server's (501 without a slot directory, "Invalid slot ID", "Invalid action", "Invalid filename", "Unable to restore slot: …"), and a file from another model or damaged by one bit is refused. 19 of 19.
  • One slot and the host prompt cache (testing/session_oracle.py → session_oracle_live). Three conversations take turns (A1 B1 C1 A2 …) through llama-server with its prompt cache on, so each turn finds another conversation's tokens in the slot: bankML must give the same text, counts and cache_n, turn by turn (modules/prompt_cache.md). Then four requests sent at once must all be answered, none mixed, each with the text it gets alone; llama-server's order among simultaneous arrivals is not deterministic, so this half checks bankML against itself. 14 of 14.
  • Logprobs (testing/logprobs_oracle.py → logprobs_oracle_live). /v1 with logprobs and top_logprobs: the same token ids, texts and bytes, every logprob the same 32-bit float, llama-server's entry rules for stop words and split UTF-8, and its refusals. Five cases are streamed, and there every chunk's delta and entry must match. 14 of 14.

5h. The q8_0 KV cache and the grammar trie (0.3.9, in progress)

  • The KV cache's kernels (oracle_ggml_b11192_q8_0_kv_kernels, forward.rs). The shipped haswell library, loaded with dlopen: quantize_row_q8_0 byte for byte and ggml_vec_dot_q8_0_q8_0 bit for bit, on random activations with outliers and on blocks of −128 and 127 where maddubs saturates. 4,000 rows quantized byte-exact, 4,000 dot products bit-exact.
  • Answers over the cache (testing/kv_oracle.py → kv_oracle_live). llama-server b11192 with --cache-type-k q8_0 --cache-type-v q8_0, and bankML with BANKML_CACHE_TYPE=q8_0: greedy and seeded, a 2,244-token prompt, a 320-token answer and a two-turn conversation; text, counts and cache_n. 6 of 6 (567 tokens). The first replay passed 2 of 6: llama.cpp rotates Q, K and V through a Hadamard transform around a quantized cache, which bankML now reproduces (fwht).
  • The grammar mask through a trie (oracle_grammar_masks, §5c). The oracle now computes every mask both ways, one candidate at a time and by the trie's walk, each against llama.cpp's mask: 196 of 196 runs, 1,645 of 1,645 masks identical to llama.cpp b11192 by each.

6. Oracles planned

  • P3, bankml's own forward pass: done. The criterion fixed here before it was built — at temperature 0, on the same prompts, an answer token-identical to llama.cpp b11192's — is met by §1d's greedy oracles (0.2.7 on), under seeded sampling too (0.2.11), and for whole conversations (0.3.0).
  • Planned (TODO.md): tool calls, each with llama-server's answers; a 4-bit (q4_0) KV cache with the same rotation; Q8_0 weights with an oracle for Qwen's own template; NEON kernels against llama.cpp's ARM build.
  • Speculative decoding (measured in 0.1.8). A draft model (Bonsai-1.7B) and n-gram speculation both left the output token-identical at temperature 0 on every run; neither was faster beyond this laptop's noise, so neither is the default (n-gram is an opt-in). The same criterion applies to any future speed-up that changes how tokens are computed.
  • The upstream Q2_0 kernel (0.2.0). upstream/test_q2_0_avx2.c holds the C drop-in to the shipped ggml_vec_dot_q2_0_q8_0, bit for bit: 200,000 of 200,000 random cases.

The rule, restated

  • An oracle is external: a compiled library, a standard's vectors, a published value, an independent implementation. A test that compares bankml with bankml is a regression test, not an oracle.
  • A comparison is exact where the quantity is exact: bits, bytes, hashes. Tolerances are used for nothing that an oracle can check exactly.
  • A speed claim is made only on code that passed every oracle in the same gate run, and only if the gain is beyond the run-to-run noise measured on the same machine.
  • Every gate run is kept (testing/results/<version>.txt), including the rejected experiments' numbers.