Loading the document…
bankml0.2.7
Verified low-bit inference for the CPU you already have — the documentation, as of 0.2.7.GitHub · Hugging Face · why bankML · the thesis · minaiml (Space)
bankml create: mindX’s persona layer, made from mindXtrain’s merged output, token-identical end to end (27 / 27, two ways) · gate record 0.3.4: mindX’s own model natively — mindx-gen39 and SmolLM2-135M-Instruct (the Llama graph in F16) and Bonsai-1.7B (tied embeddings) token-identical to llama-server b11192: whole model 800 / 800 and 840 / 840 rows bit-exact, F16 products 552,268 / 552,268 elements bit-exact against ggml, seeded sampling 40 / 40 on each, conversations 9 / 9, JSON mode 23 / 23 · JSON mode — answers under llama-server’s own json_object grammar token-identical, greedy and seeded (23 / 23 answers on the 1-bit model, 13 / 13 on the ternary), and llama.cpp’s grammar sampler matched on 1,645 / 1,645 whole-vocabulary masks · a C API, libbankml, identical to serve --native · Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on 9 / 9 turns (text, token counts, prompt-cache reuse) · own forward pass 1,064 / 1,064 rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on 40 / 40 continuations · 8,188,239,872 ternary weights and 762 / 762 dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token 0.222 s vs 2.139 sLoading the document…