# The C API (0.3.2) *Written for 0.3.2, updated for 0.3.9 (unreleased; v0.3.6 is the latest public release). The functions are the same seven since 0.3.2; what `bankml_chat` takes has grown with the engine, release by release ([CHANGELOG.md](../CHANGELOG.md)). The module page is [modules/capi.md](modules/capi.md).* bankML as a library a C, C++, Python (`ctypes`), Go (`cgo`) or Swift program links: **`libbankml.so`** and **`libbankml.a`**, declared in one hand-written header, [`capi/include/bankml.h`](../capi/include/bankml.h). The shape is llama.h's (open a model, run a chat, close it), with bankML's gate in front of it: - a model opens only after the GGUF guard plays it and its sha256 equals its `FORK.json` record. This is the same `verify` that `bankml serve` runs; - every answer carries the receipt that `bankml serve --native` gives: the model's sha256, the guard verdict, the counts, the times, and the sha256 of the answer and of the request. The crate is `capi/` (package `bankml-capi`), a second member of the workspace. Its only dependency is the bankml crate itself, by path, so **the workspace still has no external crate**. Its licence is `MIT OR Apache-2.0`, like the core. ## Build and link ```sh cargo build --release -p bankml-capi # target/release/libbankml.so and target/release/libbankml.a ``` Shared: ```sh cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app ``` Static (one binary, no `.so` to ship; the system libraries are the ones `cargo rustc -p bankml-capi --crate-type staticlib -- --print native-static-libs` lists): ```sh cc app.c -I capi/include target/release/libbankml.a -lgcc_s -lutil -lrt -lpthread -lm -ldl -o app ``` `cargo build --release` at the root builds the `bankml` binary only, as before. The root package is the default member, so the C library is built only when it is asked for. The toolchain is pinned in `rust-toolchain.toml` to **Rust 1.99.0**, because `bankml_log` is a C-variadic function defined in Rust, which was stabilized in 1.99. rustup reads that file and fetches 1.99 by itself; the machine's default toolchain is not changed. ## The API ```c const char *bankml_version(void); bankml_t *bankml_open(const char *model_path, const char *fork_json_path, uint32_t n_ctx, char **err); int bankml_chat(bankml_t *h, const char *request_json, bankml_piece_cb cb, void *user, char **result_json); void bankml_close(bankml_t *h); void bankml_free(char *s); void bankml_set_log(bankml_log_cb cb, void *user); void bankml_log(int level, const char *fmt, ...); typedef int (*bankml_piece_cb)(const char *piece, size_t len, void *user); /* return 0 to stop */ typedef void (*bankml_log_cb)(int level, const char *msg, size_t len, void *user); ``` ### `bankml_open` `bankml_open` runs the full verification before anything is mapped for inference: 1. the GGUF guard: it refuses fork-only types, legacy group-128 Q2_0 and Bonsai 2 on mainline; 2. the sha256 pin to the `FORK.json` at `fork_json_path`; 3. the file-identity check: the file must not change while it is being hashed; 4. the native check `serve --native` makes: Qwen3 in Q1_0 or Q2_0_g64, or (0.3.4) Llama in F16, tied embeddings or not, a tokenizer and a chat template bankML reproduces. On a refusal it returns NULL, with the reason in `*err` in `serve`'s words. Two examples: - `refuse: Qwen3-0.6B-Q8_0.gguf: bankML's native forward pass does not play it: weights are Q8_0 … (… O3)` - `refuse: X.gguf has no sha256 record in FORK.json: unpinned, refused` `n_ctx` is the context in tokens. 0 means 4096, `serve`'s default. ### `bankml_chat` `request_json` is the chat request that `/v1/chat/completions` takes on `serve --native`: - `messages`; - the sampling keys: `temperature`, `top_k` (1 to 128), `top_p`, `min_p`, `min_keep` and `seed`; since 0.3.6 the penalties `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`; since 0.3.7 `typical_p`, `top_n_sigma`, `xtc_probability`, `xtc_threshold`, `dynatemp_range`, `dynatemp_exponent` and the `dry_*` keys. A key that is not set takes the model's GGUF default. The request is read by the same code as `/v1` (`serve::NativeChat::parse`), so each key behaves as it does there, token-identical to llama-server b11192; - `max_tokens` (or llama-server's `n_predict`); - `stop`; - since 0.3.3, `response_format` (`{"type": "json_object"}`, or the schemas `{}` / `{"type": "object"}`), `json_schema` and `grammar` (GBNF), exactly as `serve --native` takes them ([usage.md](usage.md), "JSON mode and grammars"); since 0.3.5 any JSON schema (`response_format` `json_schema`, `json_object` with a `schema`, the top-level `json_schema`), converted into the grammar llama-server b11192 builds for it on the model's template, with the numbers read from the request's own text. In JSON mode and under a schema the callback receives the content's growth, and `result_json`'s `content` is the value as llama-server answers it (the raw text when nothing of the value came, as the server does). - since 0.3.8, `logprobs` and `top_logprobs`: `result_json` then carries `choices[0].logprobs.content` as the non-streamed `/v1` answer does. The live logprobs oracle checks `/v1`; the C API's chat oracle does not check logprobs. A sampler bankML does not reproduce is refused with a reason (`BANKML_E_REQUEST`), never approximated: `mirostat` (other than 0), a custom `samplers` order, and `top_k` outside 1 to 128. So is a JSON schema llama.cpp b11192 itself refuses (with its message; 0.3.5: every other schema is converted as llama-server converts it), a grammar llama.cpp would not parse, and (0.3.8) a prompt that does not fit the context, whose message is llama-server's 400 body. `model` and `stream` are ignored: the handle names the model, and the callback is the stream. The answer streams to `cb` as whole UTF-8 pieces, each `len` bytes long and NUL-terminated. To stop, return 0. A NULL `cb` is allowed. `*result_json` receives the object that the non-streamed `/v1/chat/completions` returns: - `choices`, `usage`, and llama-server's `timings.cache_n`; - `bankml_receipt`, made by the same code that `serve` uses (`serve::NativeChat`, `serve::completion_json` and `serve::Tally`). The receipt's `ttft_ms` is set, because the C API always streams. On an error, `*result_json` is `{"error": {"code", "message"}}`. The return code is one of these: | code | meaning | |---|---| | `BANKML_OK` (0) | answered | | `BANKML_E_ARG` (−1) | a NULL handle or request, or a request that is not UTF-8 | | `BANKML_E_REQUEST` (−2) | the request is refused: not JSON, no messages, a sampler that is not reproduced, a bad `stop`, a schema or grammar that is not taken, a prompt that does not fit the context (0.3.8) | | `BANKML_E_CHANGED` (−3) | the model file changed since it was verified. There is no answer; open it again to verify it again | | `BANKML_E_ENGINE` (−4) | the engine failed while answering | | `BANKML_E_PANIC` (−5) | a panic was caught (unwinding builds only; see below) | The handle keeps one slot, as llama-server's `-np 1` does. The next request reuses the longest common prefix of the tokens already in its KV cache. A conversation sent turn by turn therefore pays only for its new tokens, and `timings.cache_n` says how many were reused. Since 0.3.8 the slot also has llama-server's host prompt cache (`BANKML_CACHE_RAM`), so conversations that take turns on one handle get llama-server's answers ([modules/prompt_cache.md](modules/prompt_cache.md)). The engine reads its environment when the handle opens, so `BANKML_CACHE_TYPE=q8_0` (0.3.9) gives the handle a `q8_0` KV cache. Slot save and restore are on `serve` only. ### Ownership - The strings bankML returns through `char **` (`err` and `result_json`) belong to the caller. Free each one once with `bankml_free`. - `bankml_version()` returns a static string. Do not free it. - Strings passed in are only read during the call. - `bankml_close` releases the model's mapping and its KV cache. NULL is a no-op for `bankml_close` and `bankml_free`. ### Threads - **One handle has one slot, so calls on it serialize.** A second thread's `bankml_chat` waits for the first one to finish. - Different handles are independent. Each maps its model, so two handles on the same 8B file share the page cache but not their KV caches. - `bankml_close` must not race a call on the same handle. - The streaming callback runs on the calling thread. It must not close the handle. - The log sink is process-wide. It may be called from any thread that calls into bankML. ### Panics and errors Nothing unwinds into C. Every entry point checks its pointers and returns NULL or an error code with a message. Each one also runs its work inside `catch_unwind`. That catches in an unwinding build, such as `cargo test`. The release profile is built with `panic = "abort"`, so there a panic aborts the process, as a failed C `assert` does; it is never undefined behaviour. The C API's own code returns errors rather than panicking (no `unwrap` on input, NULL checked, a NUL inside a message replaced rather than refused). The engine underneath is not proven panic-free, so a panic there is a bug to report. ## The log: a C-variadic function defined in Rust ```c void bankml_set_log(bankml_log_cb cb, void *user); /* NULL: standard error */ void bankml_log(int level, const char *fmt, ...); /* BANKML_LOG_ERROR 0 · WARN 1 · INFO 2 · DEBUG 3 */ ``` `bankml_log` is defined in Rust with Rust 1.99's C-variadic function definitions: ```rust pub unsafe extern "C" fn bankml_log(level: c_int, fmt: *const c_char, mut args: ...) ``` Its formatter (`capi/src/printf.rs`) reads each argument with `args.next_arg::()`, at exactly the C type that the conversion names. The library's own messages go to the same sink: a model verified, the GPU joining or not, a refusal. The core has a small log hook for this (`bankml::log`, `bankml::set_log_sink`), which writes to standard error as before unless an embedder installs a sink. The callback receives the message, its length in bytes (it may contain a NUL if a `%c` printed one) and a terminating NUL. **Supported**, byte-identical to glibc's `snprintf` (proven by the printf oracle below): | conversion | lengths | notes | |---|---|---| | `%d %i` | none, `hh`, `h`, `l`, `ll`, `z` | `int`, `signed char`, `short`, `long`, `long long`, `ssize_t` | | `%u %x %X` | none, `hh`, `h`, `l`, `ll`, `z` | `#` gives `0x`/`0X` before a non-zero value; `+` and space are ignored, as C ignores them | | `%f %F` | none, `l` | precision (default 6), exact decimal expansion, ties to even; `inf`/`nan` (`INF`/`NAN`), a sign on `-nan` and `-0.0` | | `%c` | none | width and `-` | | `%s` | none | width, `-`, precision (reads at most that many bytes, so the string need not be NUL-terminated); NULL prints `(null)` | | `%p` | none | `0x…`, NULL prints `(nil)`; width and `-` | | `%%` | | | The flags are `-`, `+`, space, `#` and `0`. A width or a precision is a number or `*` (an `int` argument; a negative width means `-`, and a negative precision means none). **Everything else is written literally, never guessed:** - **An unknown conversion or length** (`%e %g %o %a %n %Lf %ls %jd`, …) is written `%`, and **no argument is read for it**. Its type is unknown, so no later argument can be located either: every later conversion is written `%`. A `%%` still prints `%`. - **A known conversion with flags it does not vouch for** is written `%`. Examples: `%+s`, `%05c`, `%#p`, `%#d`, `%.3c`, and a width or precision over 65,536. Its argument has a known type, so it is read and dropped, and the rest of the format goes on. - **`%n` never writes** through its argument. - A `%` at the very end is written `%`. ## The oracles (in the release gate) - **`printf oracle`** (`testing/capi/printf_oracle.c`, compiled with the system `cc`). For every case, the message that `bankml_log` delivers to a sink installed with `bankml_set_log` must equal what libc `snprintf` writes for the same format and arguments. The comparison is byte for byte, with the length, so an embedded NUL counts. The cases cover: - integers at their limits in every length; - `%f` precision and rounding (ties, `%.0f` of 0.5/1.5/2.5, `%.1100f` of the smallest subnormal, `DBL_MAX`); - `inf` and `nan` with signs; - strings, including a precision-bounded unterminated buffer and NULL; - `%c` of 0 and 255, `%p`, `%%`; - widths, precisions and `*`. Then every unsupported case must give its exact marker and read no argument it cannot type, and `%n` must leave its target untouched. The Rust unit tests (`cargo test -p bankml-capi`) run the same comparison from Rust, calling the Rust-defined variadic function, plus a typed-queue model that fails if the formatter reads a type other than the one passed. - **`capi chat oracle`** (`testing/capi/chat.c` + `capi_oracle.py --chat`). A C program embeds `libbankml` the way an application would. It has three cases: - **Ternary-Bonsai-8B.** `serve_oracle.py`'s three Savante conversations (9 turns) go through `bankml serve --native` `/v1/chat/completions`. Then the very same request bytes go through `bankml_chat` in a fresh process. These must be equal, turn by turn: the text, the streamed pieces joined, the prompt and completion counts, the prompt-cache reuse, the finish reason, and the receipt's `response_sha256`, `request_sha256` and `model_sha256`. - **Bonsai-8B Q1_0.** llama-server b11192's recorded conversations go through `bankml_chat`, and the same text and counts must come back, with `response_sha256` equal to the sha256 of llama-server's text. The gate's `serve_oracle --bankml` ties the same record to `serve --native`. - **The O4 models (0.3.4).** Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39, each against its own llama-server record, as Bonsai-8B above. - **Refusals.** Qwen3-0.6B Q8_0 is refused (Q8_0 weights, as `serve --native` refuses them), and so is a file that its `FORK.json` does not pin. The library's own log reaches the installed sink. ## Why it exists This C API is the llama.h-shaped seam for embedding: a program that today links llama.cpp's `llama.h` can link bankML instead and get the same answers (the oracles above), with the verification and the receipt built in. It is also the path to the handheld targets (P5): an Android NDK or iOS app embeds a C library, not an HTTP server. What it does not have yet is listed in [TODO.md](TODO.md#032--toolchain-and-the-c-api-done): tokenizer access, raw logits, slot save and restore, more than one handle sharing one mapping, and a stable ABI promise (semver for the C API comes with 1.0.0).