Download docs/CAPI.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 14.9 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/CAPI.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/CAPI.md
-
curl -L -o CAPI.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/CAPI.md
The C API (0.3.2)
Written for 0.3.2, updated for 0.3.9 (unreleased; v0.3.6 is the latest public release). The functions are the same
seven since 0.3.2; what bankml_chat takes has grown with the engine, release by release (CHANGELOG.md).
The module page is modules/capi.md.
bankML as a library a C, C++, Python (ctypes), Go (cgo) or Swift program links: libbankml.so and
libbankml.a, declared in one hand-written header, capi/include/bankml.h. The
shape is llama.h's (open a model, run a chat, close it), with bankML's gate in front of it:
- a model opens only after the GGUF guard plays it and its sha256 equals its
FORK.jsonrecord. This is the sameverifythatbankml serveruns; - every answer carries the receipt that
bankml serve --nativegives: the model's sha256, the guard verdict, the counts, the times, and the sha256 of the answer and of the request.
The crate is capi/ (package bankml-capi), a second member of the workspace. Its only dependency is the bankml
crate itself, by path, so the workspace still has no external crate. Its licence is MIT OR Apache-2.0, like the
core.
Build and link
cargo build --release -p bankml-capi # target/release/libbankml.so and target/release/libbankml.a
Shared:
cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app
Static (one binary, no .so to ship; the system libraries are the ones cargo rustc -p bankml-capi --crate-type staticlib -- --print native-static-libs lists):
cc app.c -I capi/include target/release/libbankml.a -lgcc_s -lutil -lrt -lpthread -lm -ldl -o app
cargo build --release at the root builds the bankml binary only, as before. The root package is the default
member, so the C library is built only when it is asked for. The toolchain is pinned in rust-toolchain.toml to
Rust 1.99.0, because bankml_log is a C-variadic function defined in Rust, which was stabilized in 1.99. rustup reads
that file and fetches 1.99 by itself; the machine's default toolchain is not changed.
The API
const char *bankml_version(void);
bankml_t *bankml_open(const char *model_path, const char *fork_json_path, uint32_t n_ctx, char **err);
int bankml_chat(bankml_t *h, const char *request_json, bankml_piece_cb cb, void *user, char **result_json);
void bankml_close(bankml_t *h);
void bankml_free(char *s);
void bankml_set_log(bankml_log_cb cb, void *user);
void bankml_log(int level, const char *fmt, ...);
typedef int (*bankml_piece_cb)(const char *piece, size_t len, void *user); /* return 0 to stop */
typedef void (*bankml_log_cb)(int level, const char *msg, size_t len, void *user);
bankml_open
bankml_open runs the full verification before anything is mapped for inference:
- the GGUF guard: it refuses fork-only types, legacy group-128 Q2_0 and Bonsai 2 on mainline;
- the sha256 pin to the
FORK.jsonatfork_json_path; - the file-identity check: the file must not change while it is being hashed;
- the native check
serve --nativemakes: Qwen3 in Q1_0 or Q2_0_g64, or (0.3.4) Llama in F16, tied embeddings or not, a tokenizer and a chat template bankML reproduces.
On a refusal it returns NULL, with the reason in *err in serve's words. Two examples:
refuse: Qwen3-0.6B-Q8_0.gguf: bankML's native forward pass does not play it: weights are Q8_0 … (… O3)refuse: X.gguf has no sha256 record in FORK.json: unpinned, refused
n_ctx is the context in tokens. 0 means 4096, serve's default.
bankml_chat
request_json is the chat request that /v1/chat/completions takes on serve --native:
messages;the sampling keys:
temperature,top_k(1 to 128),top_p,min_p,min_keepandseed; since 0.3.6 the penaltiesrepeat_penalty,repeat_last_n,presence_penalty,frequency_penalty; since 0.3.7typical_p,top_n_sigma,xtc_probability,xtc_threshold,dynatemp_range,dynatemp_exponentand thedry_*keys. A key that is not set takes the model's GGUF default. The request is read by the same code as/v1(serve::NativeChat::parse), so each key behaves as it does there, token-identical to llama-server b11192;max_tokens(or llama-server'sn_predict);stop;since 0.3.3,
response_format({"type": "json_object"}, or the schemas{}/{"type": "object"}),json_schemaandgrammar(GBNF), exactly asserve --nativetakes them (usage.md, "JSON mode and grammars"); since 0.3.5 any JSON schema (response_formatjson_schema,json_objectwith aschema, the top-leveljson_schema), converted into the grammar llama-server b11192 builds for it on the model's template, with the numbers read from the request's own text. In JSON mode and under a schema the callback receives the content's growth, andresult_json'scontentis the value as llama-server answers it (the raw text when nothing of the value came, as the server does).since 0.3.8,
logprobsandtop_logprobs:result_jsonthen carrieschoices[0].logprobs.contentas the non-streamed/v1answer does. The live logprobs oracle checks/v1; the C API's chat oracle does not check logprobs.
A sampler bankML does not reproduce is refused with a reason (BANKML_E_REQUEST), never approximated: mirostat
(other than 0), a custom samplers order, and top_k outside 1 to 128. So is a JSON schema llama.cpp b11192 itself
refuses (with its message; 0.3.5: every other schema is converted as llama-server converts it), a grammar llama.cpp
would not parse, and (0.3.8) a prompt that does not fit the context, whose message is llama-server's 400 body.
model and stream are ignored: the handle names the model, and the callback is the stream.
The answer streams to cb as whole UTF-8 pieces, each len bytes long and NUL-terminated. To stop, return 0. A NULL
cb is allowed.
*result_json receives the object that the non-streamed /v1/chat/completions returns:
choices,usage, and llama-server'stimings.cache_n;bankml_receipt, made by the same code thatserveuses (serve::NativeChat,serve::completion_jsonandserve::Tally).
The receipt's ttft_ms is set, because the C API always streams. On an error, *result_json is
{"error": {"code", "message"}}.
The return code is one of these:
| code | meaning |
|---|---|
BANKML_OK (0) |
answered |
BANKML_E_ARG (−1) |
a NULL handle or request, or a request that is not UTF-8 |
BANKML_E_REQUEST (−2) |
the request is refused: not JSON, no messages, a sampler that is not reproduced, a bad stop, a schema or grammar that is not taken, a prompt that does not fit the context (0.3.8) |
BANKML_E_CHANGED (−3) |
the model file changed since it was verified. There is no answer; open it again to verify it again |
BANKML_E_ENGINE (−4) |
the engine failed while answering |
BANKML_E_PANIC (−5) |
a panic was caught (unwinding builds only; see below) |
The handle keeps one slot, as llama-server's -np 1 does. The next request reuses the longest common prefix of the
tokens already in its KV cache. A conversation sent turn by turn therefore pays only for its new tokens, and
timings.cache_n says how many were reused. Since 0.3.8 the slot also has llama-server's host prompt cache
(BANKML_CACHE_RAM), so conversations that take turns on one handle get llama-server's answers
(modules/prompt_cache.md). The engine reads its environment when the handle opens, so
BANKML_CACHE_TYPE=q8_0 (0.3.9) gives the handle a q8_0 KV cache. Slot save and restore are on serve only.
Ownership
- The strings bankML returns through
char **(errandresult_json) belong to the caller. Free each one once withbankml_free. bankml_version()returns a static string. Do not free it.- Strings passed in are only read during the call.
bankml_closereleases the model's mapping and its KV cache. NULL is a no-op forbankml_closeandbankml_free.
Threads
- One handle has one slot, so calls on it serialize. A second thread's
bankml_chatwaits for the first one to finish. - Different handles are independent. Each maps its model, so two handles on the same 8B file share the page cache but not their KV caches.
bankml_closemust not race a call on the same handle.- The streaming callback runs on the calling thread. It must not close the handle.
- The log sink is process-wide. It may be called from any thread that calls into bankML.
Panics and errors
Nothing unwinds into C. Every entry point checks its pointers and returns NULL or an error code with a message. Each
one also runs its work inside catch_unwind. That catches in an unwinding build, such as cargo test. The release
profile is built with panic = "abort", so there a panic aborts the process, as a failed C assert does; it is never
undefined behaviour. The C API's own code returns errors rather than panicking (no unwrap on input, NULL checked, a NUL inside a
message replaced rather than refused). The engine underneath is not proven panic-free, so a panic there is a bug to
report.
The log: a C-variadic function defined in Rust
void bankml_set_log(bankml_log_cb cb, void *user); /* NULL: standard error */
void bankml_log(int level, const char *fmt, ...); /* BANKML_LOG_ERROR 0 · WARN 1 · INFO 2 · DEBUG 3 */
bankml_log is defined in Rust with Rust 1.99's C-variadic function definitions:
pub unsafe extern "C" fn bankml_log(level: c_int, fmt: *const c_char, mut args: ...)
Its formatter (capi/src/printf.rs) reads each argument with args.next_arg::<T>(), at exactly the C type that the
conversion names.
The library's own messages go to the same sink: a model verified, the GPU joining or not, a refusal. The core has a
small log hook for this (bankml::log, bankml::set_log_sink), which writes to standard error as before unless an
embedder installs a sink. The callback receives the message, its length in bytes (it may contain a NUL if a %c
printed one) and a terminating NUL.
Supported, byte-identical to glibc's snprintf (proven by the printf oracle below):
| conversion | lengths | notes |
|---|---|---|
%d %i |
none, hh, h, l, ll, z |
int, signed char, short, long, long long, ssize_t |
%u %x %X |
none, hh, h, l, ll, z |
# gives 0x/0X before a non-zero value; + and space are ignored, as C ignores them |
%f %F |
none, l |
precision (default 6), exact decimal expansion, ties to even; inf/nan (INF/NAN), a sign on -nan and -0.0 |
%c |
none | width and - |
%s |
none | width, -, precision (reads at most that many bytes, so the string need not be NUL-terminated); NULL prints (null) |
%p |
none | 0x…, NULL prints (nil); width and - |
%% |
The flags are -, +, space, # and 0. A width or a precision is a number or * (an int argument; a negative
width means -, and a negative precision means none).
Everything else is written literally, never guessed:
- An unknown conversion or length (
%e %g %o %a %n %Lf %ls %jd, …) is written%<unsupported:SPEC>, and no argument is read for it. Its type is unknown, so no later argument can be located either: every later conversion is written%<skipped:SPEC>. A%%still prints%. - A known conversion with flags it does not vouch for is written
%<unsupported:SPEC>. Examples:%+s,%05c,%#p,%#d,%.3c, and a width or precision over 65,536. Its argument has a known type, so it is read and dropped, and the rest of the format goes on. %nnever writes through its argument.- A
%at the very end is written%<incomplete:>.
The oracles (in the release gate)
printf oracle(testing/capi/printf_oracle.c, compiled with the systemcc). For every case, the message thatbankml_logdelivers to a sink installed withbankml_set_logmust equal what libcsnprintfwrites for the same format and arguments. The comparison is byte for byte, with the length, so an embedded NUL counts. The cases cover:- integers at their limits in every length;
%fprecision and rounding (ties,%.0fof 0.5/1.5/2.5,%.1100fof the smallest subnormal,DBL_MAX);infandnanwith signs;- strings, including a precision-bounded unterminated buffer and NULL;
%cof 0 and 255,%p,%%;- widths, precisions and
*.
Then every unsupported case must give its exact marker and read no argument it cannot type, and
%nmust leave its target untouched. The Rust unit tests (cargo test -p bankml-capi) run the same comparison from Rust, calling the Rust-defined variadic function, plus a typed-queue model that fails if the formatter reads a type other than the one passed.capi chat oracle(testing/capi/chat.c+capi_oracle.py --chat). A C program embedslibbankmlthe way an application would. It has three cases:- Ternary-Bonsai-8B.
serve_oracle.py's three Savante conversations (9 turns) go throughbankml serve --native/v1/chat/completions. Then the very same request bytes go throughbankml_chatin a fresh process. These must be equal, turn by turn: the text, the streamed pieces joined, the prompt and completion counts, the prompt-cache reuse, the finish reason, and the receipt'sresponse_sha256,request_sha256andmodel_sha256. - Bonsai-8B Q1_0. llama-server b11192's recorded conversations go through
bankml_chat, and the same text and counts must come back, withresponse_sha256equal to the sha256 of llama-server's text. The gate'sserve_oracle --bankmlties the same record toserve --native. - The O4 models (0.3.4). Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39, each against its own llama-server record, as Bonsai-8B above.
- Refusals. Qwen3-0.6B Q8_0 is refused (Q8_0 weights, as
serve --nativerefuses them), and so is a file that itsFORK.jsondoes not pin. The library's own log reaches the installed sink.
- Ternary-Bonsai-8B.
Why it exists
This C API is the llama.h-shaped seam for embedding: a program that today links llama.cpp's llama.h can link
bankML instead and get the same answers (the oracles above), with the verification and the receipt built in. It is
also the path to the handheld targets (P5): an Android NDK or iOS app embeds a C library, not an HTTP server. What it
does not have yet is listed in TODO.md: tokenizer access, raw logits,
slot save and restore, more than one handle sharing one mapping, and a stable ABI promise (semver for the C API comes
with 1.0.0).