|
Download docs/CAPI.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 14.9 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/CAPI.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/CAPI.md
-
curl -L -o CAPI.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/CAPI.md
14.9 kB
| # The C API (0.3.2) | |
| *Written for 0.3.2, updated for 0.3.9 (unreleased; v0.3.6 is the latest public release). The functions are the same | |
| seven since 0.3.2; what `bankml_chat` takes has grown with the engine, release by release ([CHANGELOG.md](../CHANGELOG.md)). | |
| The module page is [modules/capi.md](modules/capi.md).* | |
| bankML as a library a C, C++, Python (`ctypes`), Go (`cgo`) or Swift program links: **`libbankml.so`** and | |
| **`libbankml.a`**, declared in one hand-written header, [`capi/include/bankml.h`](../capi/include/bankml.h). The | |
| shape is llama.h's (open a model, run a chat, close it), with bankML's gate in front of it: | |
| - a model opens only after the GGUF guard plays it and its sha256 equals its `FORK.json` record. This is the same | |
| `verify` that `bankml serve` runs; | |
| - every answer carries the receipt that `bankml serve --native` gives: the model's sha256, the guard verdict, the | |
| counts, the times, and the sha256 of the answer and of the request. | |
| The crate is `capi/` (package `bankml-capi`), a second member of the workspace. Its only dependency is the bankml | |
| crate itself, by path, so **the workspace still has no external crate**. Its licence is `MIT OR Apache-2.0`, like the | |
| core. | |
| ## Build and link | |
| ```sh | |
| cargo build --release -p bankml-capi # target/release/libbankml.so and target/release/libbankml.a | |
| ``` | |
| Shared: | |
| ```sh | |
| cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app | |
| ``` | |
| Static (one binary, no `.so` to ship; the system libraries are the ones `cargo rustc -p bankml-capi --crate-type | |
| staticlib -- --print native-static-libs` lists): | |
| ```sh | |
| cc app.c -I capi/include target/release/libbankml.a -lgcc_s -lutil -lrt -lpthread -lm -ldl -o app | |
| ``` | |
| `cargo build --release` at the root builds the `bankml` binary only, as before. The root package is the default | |
| member, so the C library is built only when it is asked for. The toolchain is pinned in `rust-toolchain.toml` to | |
| **Rust 1.99.0**, because `bankml_log` is a C-variadic function defined in Rust, which was stabilized in 1.99. rustup reads | |
| that file and fetches 1.99 by itself; the machine's default toolchain is not changed. | |
| ## The API | |
| ```c | |
| const char *bankml_version(void); | |
| bankml_t *bankml_open(const char *model_path, const char *fork_json_path, uint32_t n_ctx, char **err); | |
| int bankml_chat(bankml_t *h, const char *request_json, bankml_piece_cb cb, void *user, char **result_json); | |
| void bankml_close(bankml_t *h); | |
| void bankml_free(char *s); | |
| void bankml_set_log(bankml_log_cb cb, void *user); | |
| void bankml_log(int level, const char *fmt, ...); | |
| typedef int (*bankml_piece_cb)(const char *piece, size_t len, void *user); /* return 0 to stop */ | |
| typedef void (*bankml_log_cb)(int level, const char *msg, size_t len, void *user); | |
| ``` | |
| ### `bankml_open` | |
| `bankml_open` runs the full verification before anything is mapped for inference: | |
| 1. the GGUF guard: it refuses fork-only types, legacy group-128 Q2_0 and Bonsai 2 on mainline; | |
| 2. the sha256 pin to the `FORK.json` at `fork_json_path`; | |
| 3. the file-identity check: the file must not change while it is being hashed; | |
| 4. the native check `serve --native` makes: Qwen3 in Q1_0 or Q2_0_g64, or (0.3.4) Llama in F16, tied embeddings or | |
| not, a tokenizer and a chat template bankML reproduces. | |
| On a refusal it returns NULL, with the reason in `*err` in `serve`'s words. Two examples: | |
| - `refuse: Qwen3-0.6B-Q8_0.gguf: bankML's native forward pass does not play it: weights are Q8_0 … (… O3)` | |
| - `refuse: X.gguf has no sha256 record in FORK.json: unpinned, refused` | |
| `n_ctx` is the context in tokens. 0 means 4096, `serve`'s default. | |
| ### `bankml_chat` | |
| `request_json` is the chat request that `/v1/chat/completions` takes on `serve --native`: | |
| - `messages`; | |
| - the sampling keys: `temperature`, `top_k` (1 to 128), `top_p`, `min_p`, `min_keep` and `seed`; since 0.3.6 the | |
| penalties `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`; since 0.3.7 `typical_p`, | |
| `top_n_sigma`, `xtc_probability`, `xtc_threshold`, `dynatemp_range`, `dynatemp_exponent` and the `dry_*` keys. A | |
| key that is not set takes the model's GGUF default. The request is read by the same code as `/v1` | |
| (`serve::NativeChat::parse`), so each key behaves as it does there, token-identical to llama-server b11192; | |
| - `max_tokens` (or llama-server's `n_predict`); | |
| - `stop`; | |
| - since 0.3.3, `response_format` (`{"type": "json_object"}`, or the schemas `{}` / `{"type": "object"}`), `json_schema` | |
| and `grammar` (GBNF), exactly as `serve --native` takes them ([usage.md](usage.md), "JSON mode and grammars"); since | |
| 0.3.5 any JSON schema (`response_format` `json_schema`, `json_object` with a `schema`, the top-level `json_schema`), | |
| converted into the grammar llama-server b11192 builds for it on the model's template, with the numbers read from the | |
| request's own text. In JSON mode and under a schema the callback receives the content's growth, and `result_json`'s | |
| `content` is the value as llama-server answers it (the raw text when nothing of the value came, as the server does). | |
| - since 0.3.8, `logprobs` and `top_logprobs`: `result_json` then carries `choices[0].logprobs.content` as the | |
| non-streamed `/v1` answer does. The live logprobs oracle checks `/v1`; the C API's chat oracle does not check | |
| logprobs. | |
| A sampler bankML does not reproduce is refused with a reason (`BANKML_E_REQUEST`), never approximated: `mirostat` | |
| (other than 0), a custom `samplers` order, and `top_k` outside 1 to 128. So is a JSON schema llama.cpp b11192 itself | |
| refuses (with its message; 0.3.5: every other schema is converted as llama-server converts it), a grammar llama.cpp | |
| would not parse, and (0.3.8) a prompt that does not fit the context, whose message is llama-server's 400 body. | |
| `model` and `stream` are ignored: the handle names the model, and the callback is the stream. | |
| The answer streams to `cb` as whole UTF-8 pieces, each `len` bytes long and NUL-terminated. To stop, return 0. A NULL | |
| `cb` is allowed. | |
| `*result_json` receives the object that the non-streamed `/v1/chat/completions` returns: | |
| - `choices`, `usage`, and llama-server's `timings.cache_n`; | |
| - `bankml_receipt`, made by the same code that `serve` uses (`serve::NativeChat`, `serve::completion_json` and | |
| `serve::Tally`). | |
| The receipt's `ttft_ms` is set, because the C API always streams. On an error, `*result_json` is | |
| `{"error": {"code", "message"}}`. | |
| The return code is one of these: | |
| | code | meaning | | |
| |---|---| | |
| | `BANKML_OK` (0) | answered | | |
| | `BANKML_E_ARG` (−1) | a NULL handle or request, or a request that is not UTF-8 | | |
| | `BANKML_E_REQUEST` (−2) | the request is refused: not JSON, no messages, a sampler that is not reproduced, a bad `stop`, a schema or grammar that is not taken, a prompt that does not fit the context (0.3.8) | | |
| | `BANKML_E_CHANGED` (−3) | the model file changed since it was verified. There is no answer; open it again to verify it again | | |
| | `BANKML_E_ENGINE` (−4) | the engine failed while answering | | |
| | `BANKML_E_PANIC` (−5) | a panic was caught (unwinding builds only; see below) | | |
| The handle keeps one slot, as llama-server's `-np 1` does. The next request reuses the longest common prefix of the | |
| tokens already in its KV cache. A conversation sent turn by turn therefore pays only for its new tokens, and | |
| `timings.cache_n` says how many were reused. Since 0.3.8 the slot also has llama-server's host prompt cache | |
| (`BANKML_CACHE_RAM`), so conversations that take turns on one handle get llama-server's answers | |
| ([modules/prompt_cache.md](modules/prompt_cache.md)). The engine reads its environment when the handle opens, so | |
| `BANKML_CACHE_TYPE=q8_0` (0.3.9) gives the handle a `q8_0` KV cache. Slot save and restore are on `serve` only. | |
| ### Ownership | |
| - The strings bankML returns through `char **` (`err` and `result_json`) belong to the caller. Free each one once | |
| with `bankml_free`. | |
| - `bankml_version()` returns a static string. Do not free it. | |
| - Strings passed in are only read during the call. | |
| - `bankml_close` releases the model's mapping and its KV cache. NULL is a no-op for `bankml_close` and | |
| `bankml_free`. | |
| ### Threads | |
| - **One handle has one slot, so calls on it serialize.** A second thread's `bankml_chat` waits for the first one to | |
| finish. | |
| - Different handles are independent. Each maps its model, so two handles on the same 8B file share the page cache | |
| but not their KV caches. | |
| - `bankml_close` must not race a call on the same handle. | |
| - The streaming callback runs on the calling thread. It must not close the handle. | |
| - The log sink is process-wide. It may be called from any thread that calls into bankML. | |
| ### Panics and errors | |
| Nothing unwinds into C. Every entry point checks its pointers and returns NULL or an error code with a message. Each | |
| one also runs its work inside `catch_unwind`. That catches in an unwinding build, such as `cargo test`. The release | |
| profile is built with `panic = "abort"`, so there a panic aborts the process, as a failed C `assert` does; it is never | |
| undefined behaviour. The C API's own code returns errors rather than panicking (no `unwrap` on input, NULL checked, a NUL inside a | |
| message replaced rather than refused). The engine underneath is not proven panic-free, so a panic there is a bug to | |
| report. | |
| ## The log: a C-variadic function defined in Rust | |
| ```c | |
| void bankml_set_log(bankml_log_cb cb, void *user); /* NULL: standard error */ | |
| void bankml_log(int level, const char *fmt, ...); /* BANKML_LOG_ERROR 0 · WARN 1 · INFO 2 · DEBUG 3 */ | |
| ``` | |
| `bankml_log` is defined in Rust with Rust 1.99's C-variadic function definitions: | |
| ```rust | |
| pub unsafe extern "C" fn bankml_log(level: c_int, fmt: *const c_char, mut args: ...) | |
| ``` | |
| Its formatter (`capi/src/printf.rs`) reads each argument with `args.next_arg::<T>()`, at exactly the C type that the | |
| conversion names. | |
| The library's own messages go to the same sink: a model verified, the GPU joining or not, a refusal. The core has a | |
| small log hook for this (`bankml::log`, `bankml::set_log_sink`), which writes to standard error as before unless an | |
| embedder installs a sink. The callback receives the message, its length in bytes (it may contain a NUL if a `%c` | |
| printed one) and a terminating NUL. | |
| **Supported**, byte-identical to glibc's `snprintf` (proven by the printf oracle below): | |
| | conversion | lengths | notes | | |
| |---|---|---| | |
| | `%d %i` | none, `hh`, `h`, `l`, `ll`, `z` | `int`, `signed char`, `short`, `long`, `long long`, `ssize_t` | | |
| | `%u %x %X` | none, `hh`, `h`, `l`, `ll`, `z` | `#` gives `0x`/`0X` before a non-zero value; `+` and space are ignored, as C ignores them | | |
| | `%f %F` | none, `l` | precision (default 6), exact decimal expansion, ties to even; `inf`/`nan` (`INF`/`NAN`), a sign on `-nan` and `-0.0` | | |
| | `%c` | none | width and `-` | | |
| | `%s` | none | width, `-`, precision (reads at most that many bytes, so the string need not be NUL-terminated); NULL prints `(null)` | | |
| | `%p` | none | `0x…`, NULL prints `(nil)`; width and `-` | | |
| | `%%` | | | | |
| The flags are `-`, `+`, space, `#` and `0`. A width or a precision is a number or `*` (an `int` argument; a negative | |
| width means `-`, and a negative precision means none). | |
| **Everything else is written literally, never guessed:** | |
| - **An unknown conversion or length** (`%e %g %o %a %n %Lf %ls %jd`, …) is written `%<unsupported:SPEC>`, and **no | |
| argument is read for it**. Its type is unknown, so no later argument can be located either: every later conversion | |
| is written `%<skipped:SPEC>`. A `%%` still prints `%`. | |
| - **A known conversion with flags it does not vouch for** is written `%<unsupported:SPEC>`. Examples: `%+s`, `%05c`, | |
| `%#p`, `%#d`, `%.3c`, and a width or precision over 65,536. Its argument has a known type, so it is read and | |
| dropped, and the rest of the format goes on. | |
| - **`%n` never writes** through its argument. | |
| - A `%` at the very end is written `%<incomplete:>`. | |
| ## The oracles (in the release gate) | |
| - **`printf oracle`** (`testing/capi/printf_oracle.c`, compiled with the system `cc`). For every case, the message | |
| that `bankml_log` delivers to a sink installed with `bankml_set_log` must equal what libc `snprintf` writes for the | |
| same format and arguments. The comparison is byte for byte, with the length, so an embedded NUL counts. The cases | |
| cover: | |
| - integers at their limits in every length; | |
| - `%f` precision and rounding (ties, `%.0f` of 0.5/1.5/2.5, `%.1100f` of the smallest subnormal, `DBL_MAX`); | |
| - `inf` and `nan` with signs; | |
| - strings, including a precision-bounded unterminated buffer and NULL; | |
| - `%c` of 0 and 255, `%p`, `%%`; | |
| - widths, precisions and `*`. | |
| Then every unsupported case must give its exact marker and read no argument it cannot type, and `%n` must leave its | |
| target untouched. The Rust unit tests (`cargo test -p bankml-capi`) run the same comparison from Rust, calling the | |
| Rust-defined variadic function, plus a typed-queue model that fails if the formatter reads a type other than the | |
| one passed. | |
| - **`capi chat oracle`** (`testing/capi/chat.c` + `capi_oracle.py --chat`). A C program embeds `libbankml` the way an | |
| application would. It has three cases: | |
| - **Ternary-Bonsai-8B.** `serve_oracle.py`'s three Savante conversations (9 turns) go through `bankml serve | |
| --native` `/v1/chat/completions`. Then the very same request bytes go through `bankml_chat` in a fresh process. | |
| These must be equal, turn by turn: the text, the streamed pieces joined, the prompt and completion counts, the | |
| prompt-cache reuse, the finish reason, and the receipt's `response_sha256`, `request_sha256` and `model_sha256`. | |
| - **Bonsai-8B Q1_0.** llama-server b11192's recorded conversations go through `bankml_chat`, and the same text and | |
| counts must come back, with `response_sha256` equal to the sha256 of llama-server's text. The gate's | |
| `serve_oracle --bankml` ties the same record to `serve --native`. | |
| - **The O4 models (0.3.4).** Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39, each against its own | |
| llama-server record, as Bonsai-8B above. | |
| - **Refusals.** Qwen3-0.6B Q8_0 is refused (Q8_0 weights, as `serve --native` refuses them), and so is a file that | |
| its `FORK.json` does not pin. The library's own log reaches the installed sink. | |
| ## Why it exists | |
| This C API is the llama.h-shaped seam for embedding: a program that today links llama.cpp's `llama.h` can link | |
| bankML instead and get the same answers (the oracles above), with the verification and the receipt built in. It is | |
| also the path to the handheld targets (P5): an Android NDK or iOS app embeds a C library, not an HTTP server. What it | |
| does not have yet is listed in [TODO.md](TODO.md#032--toolchain-and-the-c-api-done): tokenizer access, raw logits, | |
| slot save and restore, more than one handle sharing one mapping, and a stable ABI promise (semver for the C API comes | |
| with 1.0.0). | |