bankml / docs /CAPI.md
Gregory-L's picture
bankML: the whole source (github.com/cryptoAGI/bankml @ 12ae409) and its page, with the bankML persona; the live engine (Dockerfile, hf/start.sh) ready for Docker hardware
28c70af verified
|
Raw History Blame Contribute Delete
14.9 kB

The C API (0.3.2)

Written for 0.3.2, updated for 0.3.9 (unreleased; v0.3.6 is the latest public release). The functions are the same seven since 0.3.2; what bankml_chat takes has grown with the engine, release by release (CHANGELOG.md). The module page is modules/capi.md.

bankML as a library a C, C++, Python (ctypes), Go (cgo) or Swift program links: libbankml.so and libbankml.a, declared in one hand-written header, capi/include/bankml.h. The shape is llama.h's (open a model, run a chat, close it), with bankML's gate in front of it:

  • a model opens only after the GGUF guard plays it and its sha256 equals its FORK.json record. This is the same verify that bankml serve runs;
  • every answer carries the receipt that bankml serve --native gives: the model's sha256, the guard verdict, the counts, the times, and the sha256 of the answer and of the request.

The crate is capi/ (package bankml-capi), a second member of the workspace. Its only dependency is the bankml crate itself, by path, so the workspace still has no external crate. Its licence is MIT OR Apache-2.0, like the core.

Build and link

cargo build --release -p bankml-capi        # target/release/libbankml.so and target/release/libbankml.a

Shared:

cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app

Static (one binary, no .so to ship; the system libraries are the ones cargo rustc -p bankml-capi --crate-type staticlib -- --print native-static-libs lists):

cc app.c -I capi/include target/release/libbankml.a -lgcc_s -lutil -lrt -lpthread -lm -ldl -o app

cargo build --release at the root builds the bankml binary only, as before. The root package is the default member, so the C library is built only when it is asked for. The toolchain is pinned in rust-toolchain.toml to Rust 1.99.0, because bankml_log is a C-variadic function defined in Rust, which was stabilized in 1.99. rustup reads that file and fetches 1.99 by itself; the machine's default toolchain is not changed.

The API

const char *bankml_version(void);
bankml_t   *bankml_open(const char *model_path, const char *fork_json_path, uint32_t n_ctx, char **err);
int         bankml_chat(bankml_t *h, const char *request_json, bankml_piece_cb cb, void *user, char **result_json);
void        bankml_close(bankml_t *h);
void        bankml_free(char *s);
void        bankml_set_log(bankml_log_cb cb, void *user);
void        bankml_log(int level, const char *fmt, ...);

typedef int  (*bankml_piece_cb)(const char *piece, size_t len, void *user);          /* return 0 to stop */
typedef void (*bankml_log_cb)(int level, const char *msg, size_t len, void *user);

bankml_open

bankml_open runs the full verification before anything is mapped for inference:

  1. the GGUF guard: it refuses fork-only types, legacy group-128 Q2_0 and Bonsai 2 on mainline;
  2. the sha256 pin to the FORK.json at fork_json_path;
  3. the file-identity check: the file must not change while it is being hashed;
  4. the native check serve --native makes: Qwen3 in Q1_0 or Q2_0_g64, or (0.3.4) Llama in F16, tied embeddings or not, a tokenizer and a chat template bankML reproduces.

On a refusal it returns NULL, with the reason in *err in serve's words. Two examples:

  • refuse: Qwen3-0.6B-Q8_0.gguf: bankML's native forward pass does not play it: weights are Q8_0 … (… O3)
  • refuse: X.gguf has no sha256 record in FORK.json: unpinned, refused

n_ctx is the context in tokens. 0 means 4096, serve's default.

bankml_chat

request_json is the chat request that /v1/chat/completions takes on serve --native:

  • messages;

  • the sampling keys: temperature, top_k (1 to 128), top_p, min_p, min_keep and seed; since 0.3.6 the penalties repeat_penalty, repeat_last_n, presence_penalty, frequency_penalty; since 0.3.7 typical_p, top_n_sigma, xtc_probability, xtc_threshold, dynatemp_range, dynatemp_exponent and the dry_* keys. A key that is not set takes the model's GGUF default. The request is read by the same code as /v1 (serve::NativeChat::parse), so each key behaves as it does there, token-identical to llama-server b11192;

  • max_tokens (or llama-server's n_predict);

  • stop;

  • since 0.3.3, response_format ({"type": "json_object"}, or the schemas {} / {"type": "object"}), json_schema and grammar (GBNF), exactly as serve --native takes them (usage.md, "JSON mode and grammars"); since 0.3.5 any JSON schema (response_format json_schema, json_object with a schema, the top-level json_schema), converted into the grammar llama-server b11192 builds for it on the model's template, with the numbers read from the request's own text. In JSON mode and under a schema the callback receives the content's growth, and result_json's content is the value as llama-server answers it (the raw text when nothing of the value came, as the server does).

  • since 0.3.8, logprobs and top_logprobs: result_json then carries choices[0].logprobs.content as the non-streamed /v1 answer does. The live logprobs oracle checks /v1; the C API's chat oracle does not check logprobs.

A sampler bankML does not reproduce is refused with a reason (BANKML_E_REQUEST), never approximated: mirostat (other than 0), a custom samplers order, and top_k outside 1 to 128. So is a JSON schema llama.cpp b11192 itself refuses (with its message; 0.3.5: every other schema is converted as llama-server converts it), a grammar llama.cpp would not parse, and (0.3.8) a prompt that does not fit the context, whose message is llama-server's 400 body. model and stream are ignored: the handle names the model, and the callback is the stream.

The answer streams to cb as whole UTF-8 pieces, each len bytes long and NUL-terminated. To stop, return 0. A NULL cb is allowed.

*result_json receives the object that the non-streamed /v1/chat/completions returns:

  • choices, usage, and llama-server's timings.cache_n;
  • bankml_receipt, made by the same code that serve uses (serve::NativeChat, serve::completion_json and serve::Tally).

The receipt's ttft_ms is set, because the C API always streams. On an error, *result_json is {"error": {"code", "message"}}.

The return code is one of these:

code meaning
BANKML_OK (0) answered
BANKML_E_ARG (−1) a NULL handle or request, or a request that is not UTF-8
BANKML_E_REQUEST (−2) the request is refused: not JSON, no messages, a sampler that is not reproduced, a bad stop, a schema or grammar that is not taken, a prompt that does not fit the context (0.3.8)
BANKML_E_CHANGED (−3) the model file changed since it was verified. There is no answer; open it again to verify it again
BANKML_E_ENGINE (−4) the engine failed while answering
BANKML_E_PANIC (−5) a panic was caught (unwinding builds only; see below)

The handle keeps one slot, as llama-server's -np 1 does. The next request reuses the longest common prefix of the tokens already in its KV cache. A conversation sent turn by turn therefore pays only for its new tokens, and timings.cache_n says how many were reused. Since 0.3.8 the slot also has llama-server's host prompt cache (BANKML_CACHE_RAM), so conversations that take turns on one handle get llama-server's answers (modules/prompt_cache.md). The engine reads its environment when the handle opens, so BANKML_CACHE_TYPE=q8_0 (0.3.9) gives the handle a q8_0 KV cache. Slot save and restore are on serve only.

Ownership

  • The strings bankML returns through char ** (err and result_json) belong to the caller. Free each one once with bankml_free.
  • bankml_version() returns a static string. Do not free it.
  • Strings passed in are only read during the call.
  • bankml_close releases the model's mapping and its KV cache. NULL is a no-op for bankml_close and bankml_free.

Threads

  • One handle has one slot, so calls on it serialize. A second thread's bankml_chat waits for the first one to finish.
  • Different handles are independent. Each maps its model, so two handles on the same 8B file share the page cache but not their KV caches.
  • bankml_close must not race a call on the same handle.
  • The streaming callback runs on the calling thread. It must not close the handle.
  • The log sink is process-wide. It may be called from any thread that calls into bankML.

Panics and errors

Nothing unwinds into C. Every entry point checks its pointers and returns NULL or an error code with a message. Each one also runs its work inside catch_unwind. That catches in an unwinding build, such as cargo test. The release profile is built with panic = "abort", so there a panic aborts the process, as a failed C assert does; it is never undefined behaviour. The C API's own code returns errors rather than panicking (no unwrap on input, NULL checked, a NUL inside a message replaced rather than refused). The engine underneath is not proven panic-free, so a panic there is a bug to report.

The log: a C-variadic function defined in Rust

void bankml_set_log(bankml_log_cb cb, void *user);   /* NULL: standard error */
void bankml_log(int level, const char *fmt, ...);    /* BANKML_LOG_ERROR 0 · WARN 1 · INFO 2 · DEBUG 3 */

bankml_log is defined in Rust with Rust 1.99's C-variadic function definitions:

pub unsafe extern "C" fn bankml_log(level: c_int, fmt: *const c_char, mut args: ...)

Its formatter (capi/src/printf.rs) reads each argument with args.next_arg::<T>(), at exactly the C type that the conversion names.

The library's own messages go to the same sink: a model verified, the GPU joining or not, a refusal. The core has a small log hook for this (bankml::log, bankml::set_log_sink), which writes to standard error as before unless an embedder installs a sink. The callback receives the message, its length in bytes (it may contain a NUL if a %c printed one) and a terminating NUL.

Supported, byte-identical to glibc's snprintf (proven by the printf oracle below):

conversion lengths notes
%d %i none, hh, h, l, ll, z int, signed char, short, long, long long, ssize_t
%u %x %X none, hh, h, l, ll, z # gives 0x/0X before a non-zero value; + and space are ignored, as C ignores them
%f %F none, l precision (default 6), exact decimal expansion, ties to even; inf/nan (INF/NAN), a sign on -nan and -0.0
%c none width and -
%s none width, -, precision (reads at most that many bytes, so the string need not be NUL-terminated); NULL prints (null)
%p none 0x…, NULL prints (nil); width and -
%%

The flags are -, +, space, # and 0. A width or a precision is a number or * (an int argument; a negative width means -, and a negative precision means none).

Everything else is written literally, never guessed:

  • An unknown conversion or length (%e %g %o %a %n %Lf %ls %jd, …) is written %<unsupported:SPEC>, and no argument is read for it. Its type is unknown, so no later argument can be located either: every later conversion is written %<skipped:SPEC>. A %% still prints %.
  • A known conversion with flags it does not vouch for is written %<unsupported:SPEC>. Examples: %+s, %05c, %#p, %#d, %.3c, and a width or precision over 65,536. Its argument has a known type, so it is read and dropped, and the rest of the format goes on.
  • %n never writes through its argument.
  • A % at the very end is written %<incomplete:>.

The oracles (in the release gate)

  • printf oracle (testing/capi/printf_oracle.c, compiled with the system cc). For every case, the message that bankml_log delivers to a sink installed with bankml_set_log must equal what libc snprintf writes for the same format and arguments. The comparison is byte for byte, with the length, so an embedded NUL counts. The cases cover:

    • integers at their limits in every length;
    • %f precision and rounding (ties, %.0f of 0.5/1.5/2.5, %.1100f of the smallest subnormal, DBL_MAX);
    • inf and nan with signs;
    • strings, including a precision-bounded unterminated buffer and NULL;
    • %c of 0 and 255, %p, %%;
    • widths, precisions and *.

    Then every unsupported case must give its exact marker and read no argument it cannot type, and %n must leave its target untouched. The Rust unit tests (cargo test -p bankml-capi) run the same comparison from Rust, calling the Rust-defined variadic function, plus a typed-queue model that fails if the formatter reads a type other than the one passed.

  • capi chat oracle (testing/capi/chat.c + capi_oracle.py --chat). A C program embeds libbankml the way an application would. It has three cases:

    • Ternary-Bonsai-8B. serve_oracle.py's three Savante conversations (9 turns) go through bankml serve --native /v1/chat/completions. Then the very same request bytes go through bankml_chat in a fresh process. These must be equal, turn by turn: the text, the streamed pieces joined, the prompt and completion counts, the prompt-cache reuse, the finish reason, and the receipt's response_sha256, request_sha256 and model_sha256.
    • Bonsai-8B Q1_0. llama-server b11192's recorded conversations go through bankml_chat, and the same text and counts must come back, with response_sha256 equal to the sha256 of llama-server's text. The gate's serve_oracle --bankml ties the same record to serve --native.
    • The O4 models (0.3.4). Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39, each against its own llama-server record, as Bonsai-8B above.
    • Refusals. Qwen3-0.6B Q8_0 is refused (Q8_0 weights, as serve --native refuses them), and so is a file that its FORK.json does not pin. The library's own log reaches the installed sink.

Why it exists

This C API is the llama.h-shaped seam for embedding: a program that today links llama.cpp's llama.h can link bankML instead and get the same answers (the oracles above), with the verification and the receipt built in. It is also the path to the handheld targets (P5): an Android NDK or iOS app embeds a C library, not an HTTP server. What it does not have yet is listed in TODO.md: tokenizer access, raw logits, slot save and restore, more than one handle sharing one mapping, and a stable ABI promise (semver for the C API comes with 1.0.0).