Download docs/modules/README.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 5.25 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/README.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/README.md
-
curl -L -o README.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/README.md
bankML's source, module by module
One page per module of bankML/ (and the C API in capi/). Each says what the module does and who calls it
(Summary), its public interface with real signatures (Technical usage), the oracles and tests that prove it
(How it is verified), why it is useful and how it keeps processing efficient (Advantages and efficiency), and
what it refuses or does not do yet (Limitations), then links (See also). Where a module has history or
rationale worth keeping, it is under Design notes. Release state: 0.3.6 is the last
release; 0.3.7, 0.3.8 and 0.3.9 (in progress) are unreleased, and each page marks what they added by version
(../../CHANGELOG.md).
The code's own comments stay short and definitive β each file's header ends with Details: docs/modules/<name>.md
β and these pages are where the explanation lives (the 2026-10-06 audit moved the narrative here). The build history β
how each phase was reached, with its evidence β is in ../BUILD_HISTORY.md.
The path of a request
client βββΊ serve.rs (OpenAI /v1) ββ
βββΊ ollama.rs (Ollama /api) ββΌββΊ native.rs (registry, residency, the slot) ββΊ forward.rs (the graph, KV cache)
C API βββΊ capi (libbankml) βββββββββ β gate: bankml.rs verify = gguf.rs guard + sha256.rs pin β
β text: tokenizer.rs, chat.rs βΌ
β draw: sampler.rs, grammar.rs, schema.rs q1_0 Β· q2_0 Β· f16 Β· gpu
β keep: prompt_cache.rs (slot states in RAM) (par.rs threads)
ββ measure: metrics.rs per answer; sys.rs usage
Pages
| page | module | in one line |
|---|---|---|
| The gate | ||
| bankml.md | bankml.rs |
the crate root: verify (guard, then pin), Verified, the ggml types, the log sink |
| gguf.md | gguf.rs |
the GGUF header parser, the guard (play | refuse | need_more), Mmap |
| sha256.md | sha256.rs |
FIPS 180-4 sha256 with SHA-NI, the FORK.json pin |
| The arithmetic | ||
| q1_0.md | q1_0.rs |
ggml's 1-bit format and its product, bit-exact, AVX2 |
| q2_0.md | q2_0.rs |
ggml's ternary format, and the kernel llama.cpp lacks on x86 (~9β10Γ) |
| f16.md | f16.rs |
ggml's two F16 products, chosen by shape as ggml chooses them |
| par.md | par.rs |
the zero-dependency thread pool; bits independent of thread count |
| gpu.md | gpu/ |
Vulkan through dlopen, bankML's own SPIR-V, a verified card's share of the rows |
| sys.md | sys.rs |
machine facts and the process's usage (bankml usage, /bankml/usage): CPU, memory, GPU busy and memory, package power |
| metrics.md | metrics.rs |
bankML's own measurements of its answers: TTFT, pp and tg tokens/s, energy per token (/bankml/metrics) |
| The model | ||
| forward.md | forward.rs |
the Qwen3 and Llama graphs, the KV cache (f16, or q8_0 with llama.cpp's rotation, 0.3.9), ggml's three attention kernels |
| tokenizer.md | tokenizer.rs, unicode_letters.rs |
llama.cpp's tokenizer and pre-tokenizers |
| chat.md | chat.rs |
the chat templates, byte-identical, chosen by the template's sha |
| The answer | ||
| sampler.md | sampler.rs |
llama-server's whole default sampler chain (0.3.7), same seed same tokens |
| grammar.md | grammar.rs |
llama.cpp's GBNF engine, JSON mode and the content rule; the vocabulary trie (0.3.9) |
| schema.md | schema.rs |
json_schema_to_grammar and the chat wrapping, per template |
| Serving | ||
| native.md | native.rs |
the engine: the slot and prompt cache, the registry, one resident model |
| prompt_cache.md | prompt_cache.rs |
llama-server's host prompt cache: interleaved conversations each find their prefix again (0.3.8) |
| serve.md | serve.rs |
the gateway: OpenAI's API, receipts, the loopback rules |
| ollama.md | ollama.rs |
Ollama's API, natively |
| create.md | create.rs |
bankml create: Modelfiles as verified layers over pinned models |
| convert.md | convert.rs |
bankml convert: safetensors β GGUF, byte-identical to llama.cpp's converter |
| main.md | main.rs |
the bankml command line |
| capi.md | capi/ |
libbankml and the printf-style log; the full reference is ../CAPI.md |
| Training | ||
| train.md | train/ |
mindXtrain's author and score stages, identical to its Python |
| The console | ||
| console.md | sAGI/console.py |
bankML as itself: four tabs (Interaction, Admin, Logging, Infotags), the measured SELF block |
Install and configure: ../install.md. Use: ../usage.md. The checks: ../oracles.md. Speed: ../PERFORMANCE.md.