File size: 14,881 Bytes
28c70af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 | # The C API (0.3.2)
*Written for 0.3.2, updated for 0.3.9 (unreleased; v0.3.6 is the latest public release). The functions are the same
seven since 0.3.2; what `bankml_chat` takes has grown with the engine, release by release ([CHANGELOG.md](../CHANGELOG.md)).
The module page is [modules/capi.md](modules/capi.md).*
bankML as a library a C, C++, Python (`ctypes`), Go (`cgo`) or Swift program links: **`libbankml.so`** and
**`libbankml.a`**, declared in one hand-written header, [`capi/include/bankml.h`](../capi/include/bankml.h). The
shape is llama.h's (open a model, run a chat, close it), with bankML's gate in front of it:
- a model opens only after the GGUF guard plays it and its sha256 equals its `FORK.json` record. This is the same
`verify` that `bankml serve` runs;
- every answer carries the receipt that `bankml serve --native` gives: the model's sha256, the guard verdict, the
counts, the times, and the sha256 of the answer and of the request.
The crate is `capi/` (package `bankml-capi`), a second member of the workspace. Its only dependency is the bankml
crate itself, by path, so **the workspace still has no external crate**. Its licence is `MIT OR Apache-2.0`, like the
core.
## Build and link
```sh
cargo build --release -p bankml-capi # target/release/libbankml.so and target/release/libbankml.a
```
Shared:
```sh
cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app
```
Static (one binary, no `.so` to ship; the system libraries are the ones `cargo rustc -p bankml-capi --crate-type
staticlib -- --print native-static-libs` lists):
```sh
cc app.c -I capi/include target/release/libbankml.a -lgcc_s -lutil -lrt -lpthread -lm -ldl -o app
```
`cargo build --release` at the root builds the `bankml` binary only, as before. The root package is the default
member, so the C library is built only when it is asked for. The toolchain is pinned in `rust-toolchain.toml` to
**Rust 1.99.0**, because `bankml_log` is a C-variadic function defined in Rust, which was stabilized in 1.99. rustup reads
that file and fetches 1.99 by itself; the machine's default toolchain is not changed.
## The API
```c
const char *bankml_version(void);
bankml_t *bankml_open(const char *model_path, const char *fork_json_path, uint32_t n_ctx, char **err);
int bankml_chat(bankml_t *h, const char *request_json, bankml_piece_cb cb, void *user, char **result_json);
void bankml_close(bankml_t *h);
void bankml_free(char *s);
void bankml_set_log(bankml_log_cb cb, void *user);
void bankml_log(int level, const char *fmt, ...);
typedef int (*bankml_piece_cb)(const char *piece, size_t len, void *user); /* return 0 to stop */
typedef void (*bankml_log_cb)(int level, const char *msg, size_t len, void *user);
```
### `bankml_open`
`bankml_open` runs the full verification before anything is mapped for inference:
1. the GGUF guard: it refuses fork-only types, legacy group-128 Q2_0 and Bonsai 2 on mainline;
2. the sha256 pin to the `FORK.json` at `fork_json_path`;
3. the file-identity check: the file must not change while it is being hashed;
4. the native check `serve --native` makes: Qwen3 in Q1_0 or Q2_0_g64, or (0.3.4) Llama in F16, tied embeddings or
not, a tokenizer and a chat template bankML reproduces.
On a refusal it returns NULL, with the reason in `*err` in `serve`'s words. Two examples:
- `refuse: Qwen3-0.6B-Q8_0.gguf: bankML's native forward pass does not play it: weights are Q8_0 … (… O3)`
- `refuse: X.gguf has no sha256 record in FORK.json: unpinned, refused`
`n_ctx` is the context in tokens. 0 means 4096, `serve`'s default.
### `bankml_chat`
`request_json` is the chat request that `/v1/chat/completions` takes on `serve --native`:
- `messages`;
- the sampling keys: `temperature`, `top_k` (1 to 128), `top_p`, `min_p`, `min_keep` and `seed`; since 0.3.6 the
penalties `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`; since 0.3.7 `typical_p`,
`top_n_sigma`, `xtc_probability`, `xtc_threshold`, `dynatemp_range`, `dynatemp_exponent` and the `dry_*` keys. A
key that is not set takes the model's GGUF default. The request is read by the same code as `/v1`
(`serve::NativeChat::parse`), so each key behaves as it does there, token-identical to llama-server b11192;
- `max_tokens` (or llama-server's `n_predict`);
- `stop`;
- since 0.3.3, `response_format` (`{"type": "json_object"}`, or the schemas `{}` / `{"type": "object"}`), `json_schema`
and `grammar` (GBNF), exactly as `serve --native` takes them ([usage.md](usage.md), "JSON mode and grammars"); since
0.3.5 any JSON schema (`response_format` `json_schema`, `json_object` with a `schema`, the top-level `json_schema`),
converted into the grammar llama-server b11192 builds for it on the model's template, with the numbers read from the
request's own text. In JSON mode and under a schema the callback receives the content's growth, and `result_json`'s
`content` is the value as llama-server answers it (the raw text when nothing of the value came, as the server does).
- since 0.3.8, `logprobs` and `top_logprobs`: `result_json` then carries `choices[0].logprobs.content` as the
non-streamed `/v1` answer does. The live logprobs oracle checks `/v1`; the C API's chat oracle does not check
logprobs.
A sampler bankML does not reproduce is refused with a reason (`BANKML_E_REQUEST`), never approximated: `mirostat`
(other than 0), a custom `samplers` order, and `top_k` outside 1 to 128. So is a JSON schema llama.cpp b11192 itself
refuses (with its message; 0.3.5: every other schema is converted as llama-server converts it), a grammar llama.cpp
would not parse, and (0.3.8) a prompt that does not fit the context, whose message is llama-server's 400 body.
`model` and `stream` are ignored: the handle names the model, and the callback is the stream.
The answer streams to `cb` as whole UTF-8 pieces, each `len` bytes long and NUL-terminated. To stop, return 0. A NULL
`cb` is allowed.
`*result_json` receives the object that the non-streamed `/v1/chat/completions` returns:
- `choices`, `usage`, and llama-server's `timings.cache_n`;
- `bankml_receipt`, made by the same code that `serve` uses (`serve::NativeChat`, `serve::completion_json` and
`serve::Tally`).
The receipt's `ttft_ms` is set, because the C API always streams. On an error, `*result_json` is
`{"error": {"code", "message"}}`.
The return code is one of these:
| code | meaning |
|---|---|
| `BANKML_OK` (0) | answered |
| `BANKML_E_ARG` (−1) | a NULL handle or request, or a request that is not UTF-8 |
| `BANKML_E_REQUEST` (−2) | the request is refused: not JSON, no messages, a sampler that is not reproduced, a bad `stop`, a schema or grammar that is not taken, a prompt that does not fit the context (0.3.8) |
| `BANKML_E_CHANGED` (−3) | the model file changed since it was verified. There is no answer; open it again to verify it again |
| `BANKML_E_ENGINE` (−4) | the engine failed while answering |
| `BANKML_E_PANIC` (−5) | a panic was caught (unwinding builds only; see below) |
The handle keeps one slot, as llama-server's `-np 1` does. The next request reuses the longest common prefix of the
tokens already in its KV cache. A conversation sent turn by turn therefore pays only for its new tokens, and
`timings.cache_n` says how many were reused. Since 0.3.8 the slot also has llama-server's host prompt cache
(`BANKML_CACHE_RAM`), so conversations that take turns on one handle get llama-server's answers
([modules/prompt_cache.md](modules/prompt_cache.md)). The engine reads its environment when the handle opens, so
`BANKML_CACHE_TYPE=q8_0` (0.3.9) gives the handle a `q8_0` KV cache. Slot save and restore are on `serve` only.
### Ownership
- The strings bankML returns through `char **` (`err` and `result_json`) belong to the caller. Free each one once
with `bankml_free`.
- `bankml_version()` returns a static string. Do not free it.
- Strings passed in are only read during the call.
- `bankml_close` releases the model's mapping and its KV cache. NULL is a no-op for `bankml_close` and
`bankml_free`.
### Threads
- **One handle has one slot, so calls on it serialize.** A second thread's `bankml_chat` waits for the first one to
finish.
- Different handles are independent. Each maps its model, so two handles on the same 8B file share the page cache
but not their KV caches.
- `bankml_close` must not race a call on the same handle.
- The streaming callback runs on the calling thread. It must not close the handle.
- The log sink is process-wide. It may be called from any thread that calls into bankML.
### Panics and errors
Nothing unwinds into C. Every entry point checks its pointers and returns NULL or an error code with a message. Each
one also runs its work inside `catch_unwind`. That catches in an unwinding build, such as `cargo test`. The release
profile is built with `panic = "abort"`, so there a panic aborts the process, as a failed C `assert` does; it is never
undefined behaviour. The C API's own code returns errors rather than panicking (no `unwrap` on input, NULL checked, a NUL inside a
message replaced rather than refused). The engine underneath is not proven panic-free, so a panic there is a bug to
report.
## The log: a C-variadic function defined in Rust
```c
void bankml_set_log(bankml_log_cb cb, void *user); /* NULL: standard error */
void bankml_log(int level, const char *fmt, ...); /* BANKML_LOG_ERROR 0 · WARN 1 · INFO 2 · DEBUG 3 */
```
`bankml_log` is defined in Rust with Rust 1.99's C-variadic function definitions:
```rust
pub unsafe extern "C" fn bankml_log(level: c_int, fmt: *const c_char, mut args: ...)
```
Its formatter (`capi/src/printf.rs`) reads each argument with `args.next_arg::<T>()`, at exactly the C type that the
conversion names.
The library's own messages go to the same sink: a model verified, the GPU joining or not, a refusal. The core has a
small log hook for this (`bankml::log`, `bankml::set_log_sink`), which writes to standard error as before unless an
embedder installs a sink. The callback receives the message, its length in bytes (it may contain a NUL if a `%c`
printed one) and a terminating NUL.
**Supported**, byte-identical to glibc's `snprintf` (proven by the printf oracle below):
| conversion | lengths | notes |
|---|---|---|
| `%d %i` | none, `hh`, `h`, `l`, `ll`, `z` | `int`, `signed char`, `short`, `long`, `long long`, `ssize_t` |
| `%u %x %X` | none, `hh`, `h`, `l`, `ll`, `z` | `#` gives `0x`/`0X` before a non-zero value; `+` and space are ignored, as C ignores them |
| `%f %F` | none, `l` | precision (default 6), exact decimal expansion, ties to even; `inf`/`nan` (`INF`/`NAN`), a sign on `-nan` and `-0.0` |
| `%c` | none | width and `-` |
| `%s` | none | width, `-`, precision (reads at most that many bytes, so the string need not be NUL-terminated); NULL prints `(null)` |
| `%p` | none | `0x…`, NULL prints `(nil)`; width and `-` |
| `%%` | | |
The flags are `-`, `+`, space, `#` and `0`. A width or a precision is a number or `*` (an `int` argument; a negative
width means `-`, and a negative precision means none).
**Everything else is written literally, never guessed:**
- **An unknown conversion or length** (`%e %g %o %a %n %Lf %ls %jd`, …) is written `%<unsupported:SPEC>`, and **no
argument is read for it**. Its type is unknown, so no later argument can be located either: every later conversion
is written `%<skipped:SPEC>`. A `%%` still prints `%`.
- **A known conversion with flags it does not vouch for** is written `%<unsupported:SPEC>`. Examples: `%+s`, `%05c`,
`%#p`, `%#d`, `%.3c`, and a width or precision over 65,536. Its argument has a known type, so it is read and
dropped, and the rest of the format goes on.
- **`%n` never writes** through its argument.
- A `%` at the very end is written `%<incomplete:>`.
## The oracles (in the release gate)
- **`printf oracle`** (`testing/capi/printf_oracle.c`, compiled with the system `cc`). For every case, the message
that `bankml_log` delivers to a sink installed with `bankml_set_log` must equal what libc `snprintf` writes for the
same format and arguments. The comparison is byte for byte, with the length, so an embedded NUL counts. The cases
cover:
- integers at their limits in every length;
- `%f` precision and rounding (ties, `%.0f` of 0.5/1.5/2.5, `%.1100f` of the smallest subnormal, `DBL_MAX`);
- `inf` and `nan` with signs;
- strings, including a precision-bounded unterminated buffer and NULL;
- `%c` of 0 and 255, `%p`, `%%`;
- widths, precisions and `*`.
Then every unsupported case must give its exact marker and read no argument it cannot type, and `%n` must leave its
target untouched. The Rust unit tests (`cargo test -p bankml-capi`) run the same comparison from Rust, calling the
Rust-defined variadic function, plus a typed-queue model that fails if the formatter reads a type other than the
one passed.
- **`capi chat oracle`** (`testing/capi/chat.c` + `capi_oracle.py --chat`). A C program embeds `libbankml` the way an
application would. It has three cases:
- **Ternary-Bonsai-8B.** `serve_oracle.py`'s three Savante conversations (9 turns) go through `bankml serve
--native` `/v1/chat/completions`. Then the very same request bytes go through `bankml_chat` in a fresh process.
These must be equal, turn by turn: the text, the streamed pieces joined, the prompt and completion counts, the
prompt-cache reuse, the finish reason, and the receipt's `response_sha256`, `request_sha256` and `model_sha256`.
- **Bonsai-8B Q1_0.** llama-server b11192's recorded conversations go through `bankml_chat`, and the same text and
counts must come back, with `response_sha256` equal to the sha256 of llama-server's text. The gate's
`serve_oracle --bankml` ties the same record to `serve --native`.
- **The O4 models (0.3.4).** Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39, each against its own
llama-server record, as Bonsai-8B above.
- **Refusals.** Qwen3-0.6B Q8_0 is refused (Q8_0 weights, as `serve --native` refuses them), and so is a file that
its `FORK.json` does not pin. The library's own log reaches the installed sink.
## Why it exists
This C API is the llama.h-shaped seam for embedding: a program that today links llama.cpp's `llama.h` can link
bankML instead and get the same answers (the oracles above), with the verification and the receipt built in. It is
also the path to the handheld targets (P5): an Android NDK or iOS app embeds a C library, not an HTTP server. What it
does not have yet is listed in [TODO.md](TODO.md#032--toolchain-and-the-c-api-done): tokenizer access, raw logits,
slot save and restore, more than one handle sharing one mapping, and a stable ABI promise (semver for the C API comes
with 1.0.0).
|