Download docs/modules/gpu.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 21.1 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/gpu.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/gpu.md
-
curl -L -o gpu.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/gpu.md
bankML/gpu/ β the video-card component: find GPUs, verify them, and give a verified card a share of the 1-bit matmuls
Summary
bankML/gpu/ finds every GPU on the machine, describes it, and puts a card to work only after it has proved that it
gives the CPU kernels' bits. It is a registry of backends. Each backend is a module with one discovery function, and
nothing is linked at build time: the Vulkan backend opens its driver library at run time with dlopen, so bankML builds
and runs everywhere, and a machine without the library simply has no devices from that backend.
A card must be a real GPU (integrated or discrete) with a compute queue. Software renderers such as Mesa's llvmpipe run
on the CPU and are refused. Before a card computes anything it must pass the on-card oracle, kernels::verify_q1_0
(the oracle rule of TECHNICAL.md Β§III.4: the same bits first, then speed).
Inside the forward pass a verified card takes the first, calibrated share of a 1-bit (Q1_0) matrix's rows while
the CPU pool computes the rest β for each matrix shape where it was measured to pay, within the limit
BANKML_GPU_LIMIT sets (0.3.7). Each output element stays one device's exact dot product, so a GPU changes the speed,
never the tokens.
The component is reached three ways: the CLI (bankml gpu, bankml gpu --remote, bankml gpu --verify), the forward
pass (Weights::open opens a gpu::worker::Worker for Q1_0 models), and the tests
(gpu_q1_0_mat_vec_bit_exact). Remote cards, the GPUs Hugging Face rents, are a separate kind: listed on request,
never selected or started.
Technical usage
CLI
bankml gpu # every video card found (Vulkan, merged with /sys/class/drm) and which bankML will use, as JSON
bankml gpu --remote # the same, plus the GPUs Hugging Face rents (listed only)
bankml gpu --verify # the bit-exact kernel oracle on each selected card; exit code 2 if any card fails
bankml gpu prints report_json(remote): a devices array, the selected cards, the notes (why a backend found
nothing), and a kernels field (Q1_0 on a verified card; Q2_0 and F16 on the CPU). With --verify, each selected card prints one line per kernel and shape, or
NOT VERIFIED β <reason>; bankml will not use this card. With no usable card it prints
bankml gpu --verify: no usable GPU found; bankml runs on the CPU.
Environment
| variable | effect |
|---|---|
BANKML_GPU=off (also none, cpu) |
turns the component off: nothing is selected, the worker is not opened |
BANKML_GPU=0,2 |
picks backend indices (the index that bankml gpu reports) |
BANKML_GPU_SHARE=0.3 |
overrides the calibrated share of each 1-bit matrix's rows; clamped to 0β0.95 |
BANKML_GPU_LIMIT=0.8 |
the share of the card's memory and time bankML may use (0.3.7; clamped 0.05β1, default 0.8); see below |
mod.rs β the registry
pub type Discover = fn() -> Result<Vec<Device>, String>;
pub const BACKENDS: &[(&str, Discover)] = &[("vulkan", vulkan::devices)];
pub const REMOTE_BACKENDS: &[(&str, Discover)] = &[("huggingface", hf::devices)];
pub fn discover() -> (Vec<Device>, Vec<String>)
pub fn selected(devs: &[Device]) -> Vec<Device>
pub fn report_json(remote: bool) -> String
Deviceholds the backend's view (name, vendor and device ids,Kind, API version,device_local_bytes,host_visible_device_local,compute_queues) pluscountandusd_per_hourfor rented hardware, and an optionalSysfsrecord.KindisIntegrated,Discrete,Virtual,Cpu(a software renderer, never used),Remote(rented, never started automatically) orOther.Device::usable()is true only forIntegratedorDiscretewith at least one compute queue.discover()asks every local backend and attaches the kernel's view of the card from/sys/class/drm/card*/device: driver, VRAM and GTT totals, PCI address, NUMA node. A card is matched by vendor and device id, so it is described by what the driver says, not only by what the API reports.device_local_bytesis the sum of the device-local heaps, the memory the device addresses fastest.host_visible_device_localis true when some device-local memory is also host-visible (integrated memory, resizable BAR), so weights need no staging copy.selected()appliesBANKML_GPU, keeps the usable devices, and orders them discrete first, then by device-local memory, largest first.- A new backend (CUDA, ROCm, Metal, β¦) is a new module and one line in
BACKENDS. - A remote card is never selected automatically, because it costs money.
vulkan.rs β Vulkan through dlopen
pub fn devices() -> Result<Vec<Device>, String>
pub(super) fn instance() -> Result<(GetInstanceProcAddr, *mut c_void), String>
One API covers every GPU a Vulkan driver exposes: AMD (RADV, AMDVLK), NVIDIA, Intel (ANV), Arm, Qualcomm and Apple
(MoltenVK). The loader is opened at run time: libvulkan.so.1, libvulkan.so, libvulkan.1.dylib, then
libMoltenVK.dylib.
Every entry point comes from vkGetInstanceProcAddr. The few C structs bankML needs are declared in the file from the
Vulkan 1.x headers. devices() creates an instance, enumerates the physical devices, reads their properties, memory
heaps and queue families, and destroys the instance. With no loader it returns
no Vulkan loader (libvulkan.so.1) on this machine, and bankML carries on with the CPU. instance() creates the
instance compute uses and never destroys it; it lives for the rest of the process. The device name is read from the
first 276 bytes of VkPhysicalDeviceProperties, written into a 4 KiB buffer larger than the whole struct.
spirv.rs β bankML's own SPIR-V assembler
bankML writes its compute shaders as SPIR-V words itself, so no shader compiler is needed at build or run time (the
development machine has none, and bankML takes no crates).
Module keeps the sections apart and joins them in the order the spec requires (Module::words()). It covers only
what compute kernels use: scalar and vector types, storage buffers, push constants, the invocation ids, workgroup
memory and a control barrier, integer and float arithmetic, GLSL.std.450 Fma, and structured control flow
(selections and loops). Steps the CPU kernels fuse are written as explicit Fma.
pub fn op(&mut self, opcode: u16, result_ty: u32, operands: &[u32]) -> u32
pub fn two_sum(&mut self, f32t: u32, a: u32, b: u32) -> (u32, u32)
pub fn fma_exact(&mut self, tys: (u32, u32, u32), a: u32, b: u32, c: u32, c_mask: u32, c1: u32, c0: u32, cf0: u32) -> u32
opdecorates everyOpFMul,OpFAdd,OpFSuband extended instructionNoContraction, so a driver cannot fuse what the CPU kernels keep separate.fma_exactbuilds a correctly roundedfma(a, b, c)from plain multiplies and adds, whatever the driver does withFma.amust carry at most 12 significant bits (an f16 value does).bis split into two 12-bit halves so both partial products are exact, Knuth's TwoSum keeps every rounding error, and Boldo and Melquiond'sRN(th + RO(tl + ul))(rounding to odd, 2008) gives the FMA's single rounding.- Why it exists: RADV on the Radeon Vega 3 does not fuse
Fma. On real layer weights the 0.2.13 kernels differed from the CPU by about one ulp on 70 % of rows, and decorating theFmaNoContractiondid not change it (CHANGELOG 0.2.14).
kernels.rs β the Q1_0 kernels
pub const Q1_0_BINDINGS: u32 = 5;
pub const LOCAL_SIZE: u32 = 64;
pub fn q1_0_mat_vec() -> Vec<u32>
pub fn q1_0_mat_vec8() -> Vec<u32>
pub fn verify_q1_0(gpu: &super::compute::Gpu) -> Result<Vec<String>, String>
pub fn pack_q1_0(w: &[u8], rows: usize, n: usize) -> (Vec<f32>, Vec<u32>)
pub fn pack_act(a: &crate::q1_0::Q8Act) -> (Vec<f32>, Vec<i32>)
- Both kernels follow
q1_0::vec_dot_refstep by step: per 128-weight block, four 32-element sub-blocks; eight lanes of exact integer sums (ggml's i8 negation wraps β128 to β128);ab = d1Β·son the first sub-block andfma(d1, s, ab)after; the outerfma(d0, ab, acc)throughfma_exact; then the horizontal sum(a0+a4 + a2+a6) + (a1+a5 + a3+a7). q1_0_mat_vecuses one invocation per output row (dispatchceil(rows / 64)workgroups).q1_0_mat_vec8uses eight per row, one per accumulation lane, which is valid because each lane's chain over the blocks is independent in the CPU kernel too; a workgroup of 64 covers 8 rows. The lane sums meet in workgroup memory and lane 0 adds them in the CPU's order, giving the same bits asq1_0_mat_vecwith eight times the parallelism. Dispatchceil(rows Β· 8 / 64)workgroups.- Bindings: weight scales (f32), weight bits (u32, four per block), activation scales (f32 per 32), activation quants (i32 words of four i8), output (f32). Push constants: rows, blocks per row.
pack_q1_0repacks weights once (scales f16βf32, exact; bits as u32 words).pack_actrepacks the q8_0 activation per call. The numbers are unchanged, only aligned.
compute.rs β Vulkan compute
impl Gpu {
pub fn open(index: usize) -> Result<Gpu, String>
pub fn buffer(&self, bytes: usize) -> Result<Buffer, String>
pub fn upload<T: Copy>(&self, data: &[T]) -> Result<Buffer, String>
pub fn write<T: Copy>(&self, b: &Buffer, data: &[T])
pub fn read_f32(&self, b: &Buffer, n: usize) -> Vec<f32>
pub fn allocated(&self) -> u64 // bytes bankML's buffers hold (0.3.7, for the limiter)
pub fn free(&self, b: Buffer)
pub fn pipeline(&self, spirv: &[u32], bindings: u32, push_bytes: u32) -> Result<Pipeline, String>
pub fn run(&self, p: &Pipeline, bufs: &[&Buffer], push: &[u32], groups: u32) -> Result<(), String>
pub fn submit(&self, p: &Pipeline, bufs: &[&Buffer], push: &[u32], groups: u32) -> Result<(), String>
pub fn wait(&self) -> Result<(), String>
}
A device with one compute queue, host-visible coherent buffers mapped for their whole life (device-local and
host-visible memory when the card has it, as an integrated GPU does), pipelines from bankML's own SPIR-V, and a
dispatch that waits on a fence. run is submit then wait. One submission is in flight at a time (one command
buffer, one fence, one descriptor set per pipeline), so wait must follow each submit before the next; wait
without a pending submit blocks indefinitely. Every Vulkan call's result is checked.
On a card without device-local host-visible memory, buffers live in host-visible memory the CPU can write. Gpu is
Send but not Sync: one thread drives it at a time (the forward pass holds the worker behind a Mutex). Nothing is
destroyed on drop: a Buffer is released only by Gpu::free, pipelines are never destroyed, and dropping a Gpu
waits for the device to go idle and leaves the device objects for the driver to reclaim at process exit.
worker.rs β the GPU beside the CPU's threads
impl Worker {
pub fn open(pool: &Pool) -> Result<Option<Worker>, String>
pub fn begin(&mut self, name: &str, w: &[u8], rows: usize, a: &Q8Act) -> Result<usize, String>
pub fn finish(&mut self, out: &mut [f32]) -> Result<(), String>
}
opentakes the first selected card, runsverify_q1_0on it, and builds theq1_0_mat_vec8pipeline. It returnsNonewhen there is no usable card or the share is below 0.02.- The share comes from a calibration: the same 4096Γ4096 product timed on the card and on the CPU pool (best of
three each),
share = card rate / (card rate + CPU rate).BANKML_GPU_SHAREoverrides it. - In
Weights::mv, for eachQ1_0product,beginstarts the card on the firstshareof the rows and returns how many it took; the CPU pool computes the rest withq1_0::mat_vec_par;finishwaits and copies the card's rows in.finishis called after everybegin, even when the card took no rows, because it also records the shape's timing. - Only the card's share of each matrix is copied to it, repacked exactly, on first use. Activations wider than 16,384 are left to the CPU.
- A failure to open or verify is logged (
bankml: GPU not used: β¦) and the CPU path runs alone.
hf.rs β Hugging Face's rented GPUs
pub const HARDWARE_URL: &str = "https://huggingface.co/api/jobs/hardware";
pub fn devices() -> Result<Vec<Device>, String>
pub fn parse(text: &str) -> Result<Vec<Device>, String>
Reads the public Jobs hardware list (no token) through the system curl, because bankML has no TLS stack of its own.
Each GPU flavor becomes one Kind::Remote device with its card count, memory and price per hour (the API's
per-minute price Γ 60). A multi-card flavor such as h200x8 is one device with count 8. Nothing is started, rented
or paid for. Using a rented card means running bankML there as a Hugging Face Job, where the Vulkan backend finds it
like any local card. See huggingface.md.
The flavors range from an Nvidia T4 to 8Γ H200. Hugging Face provisions them on the large clouds (its Inference
Endpoints name AWS, Azure and GCP); NVIDIA agreed on 2026-09-02 to acquire Hugging Face. curl is used as the
installer's downloads use it; without curl or the network the backend reports why and bankML carries on with local
cards. A price per hour is reported only when the API's price is per minute.
The limiter and per-shape calibration (0.3.7)
pub fn limit() -> f64 // BANKML_GPU_LIMIT, clamped 0.05β1, default 0.8
pub fn status() -> Option<Status> // None when no card works in this process
pub fn status_json() -> String // GET /bankml/usage's "gpu_limiter"
BANKML_GPU_LIMIT= L: bankML's buffers stay within L of the heap they come from (theGpu.heap_bytesfield, againstGpu::allocated()), and on an integrated card within L of the RAM they could use (what they hold plus what the system has available). After a dispatch ofdthe card restsdΒ·(1βL)/L. A matrix that does not fit, or a product arriving while the card rests, runs whole on the CPU β never a wait, the same bits.- Per-shape calibration: each matrix shape's first products alternate between the calibrated share and the CPU
alone, timed
beginβfinishin the real pipeline (the CPU's rows included). After a warm-up (which includes the upload) and 6 timings each (TRIALS) the faster median is kept, and a shape decided for the CPU frees its buffers. A product the card skipped because it was resting or out of memory is not counted as a timing of the card. The bits are the same either way. - Why per shape: the one-off calibration measures a 4096Γ4096 product in isolation. In the forward pass a smaller matrix may not repay the submit-and-wait, and on an integrated card (whose heap is system RAM) the card and the CPU share the memory bandwidth that decode is bound by.
status_jsonreportscard,limit,share,allocated_bytes,heap_bytes,integrated,busy,products_on_card,products_on_cpu_resting,products_on_cpu_memory,shapes_on_card,shapes_on_cpuandshapes_tuning. The console's Admin tab charts it (console.md).- Measured (CHANGELOG 0.3.7; Vega 3, Bonsai-1.7B, interleaved A/B, 4 rounds, the same answer every time): with one global share the card cost decode 18 % (10.40 β 8.51β8.58 tok/s); per shape, 10.36 tok/s off against 10.30β10.38 on β neutral β with ~13 MB on the card instead of 66 MB.
How it is verified
gpu_q1_0_mat_vec_bit_exact(kernels.rs, ignored by default: needs a Vulkan GPU; in the release gate). Both kernels on every usable card against bankML's CPU kernel, which is bit-exact against ggml, on random weights and activations with β128 quants. Shapes 64Γ128 to 12288Γ4096. On the Vega 3, every row of every shape is bit-exact (oracles.md, 0.2.13).bankml gpu --verifyrunsverify_q1_0: the same comparison, plus a layer-shaped regime (scales and magnitudes that vary per block, so products are inexact in f32). With the driver'sFmaput back, it refuses the Vega 3 ("row 1 differs β¦ bankml will not use this card"). The worker runs the same check before it takes any work.- The forward pass with the card working. Every token oracle runs through
Weights::mv, so with a verified card present they all run with the GPU computing its share. In 0.2.14 the whole 1-bit model passed at 1,064 of 1,064 rows with the Vega 3 on 26 % of every matrix, and the greedy and sampling checks went through the same path. - Unit tests:
software_renderers_are_never_selected_and_discrete_cards_come_first,discovery_never_panics(mod.rs),a_multi_card_flavor_is_one_device_with_its_count(hf.rs),header_and_string_words(spirv.rs).
Advantages and efficiency
- The same bits, then speed. A card is trusted only after it reproduces the CPU kernel's bits on the card, and it works on whole rows, so no element is split between devices. The answer and the receipt do not depend on whether a GPU was present.
- No SDK, no crate, no shader compiler. Vulkan is opened with
dlopenand every entry point comes fromvkGetInstanceProcAddr; shaders are assembled in Rust byspirv.rs. The same binary runs on a machine with or without a GPU. One API covers AMD, NVIDIA, Intel, Arm and Qualcomm, and a rented NVIDIA card needs no CUDA toolkit (huggingface.md). - The card and the CPU finish together. The calibrated share gives the card the fraction of rows that matches its
rate, and the CPU pool computes the rest while the card works (
submitreturns at once;waitis called after the CPU part). - Exact FMA only where needed. The inner
fma(d1, s, ab)stays a plainFma, becaused1Β·sis exact in f32; only the outerfma(d0, ab, acc)usesfma_exact(CHANGELOG 0.2.14). - Weights copied once, and only the card's share. On an integrated GPU, buffers are allocated in device-local, host-visible memory when the card offers it.
- Measured. On the Vega 3 the eight-lane kernel matches one CPU thread, 1.7β2.0 ms per 4096Γ4096 (CHANGELOG 0.2.13). The calibrated share is 26β35 % there. Decode is unchanged within noise: 1.93β2.00 tokens/s with the card against 1.96β1.97 without (CHANGELOG 0.2.14).
- Rust practice visible in the code. No external crate (the manifest's
[dependencies]is empty).unsafeis confined to the FFI filesvulkan.rsandcompute.rs, withSAFETYcomments, behind safe methods onGpu; the registry, kernels, assembler, worker and Hugging Face backend have none. Errors areResult<_, String>with the failing call named, and a failure falls back to the CPU instead of stopping. Verification fails closed. The toolchain is pinned to Rust 1.99.0 inrust-toolchain.toml. - Where optimization goes next (TODO.md, 0.5.0): batched submissions (Q/K/V and gate/up in one
command buffer, one wait per group, persistent descriptor sets), the Q2_0 (ternary) GPU kernel in
--verify, several cards each taking a share of every matrix's rows, and a discrete-card measurement (a rented T4 or L4 run, only with the owner's approval).
Limitations
- Speed-neutral on the APU. On the Vega 3 the card is worth about one CPU core, and each of the 253 matrix calls per token pays one submit and wait, which cancels the gain (CHANGELOG 0.2.14). Per-shape calibration (0.3.7) keeps the card from costing decode, but does not make it faster there. One submission is in flight at a time.
Q1_0only. The worker is opened only for 1-bit models; there is no Q2_0 (ternary) or F16 GPU kernel yet.- One card. The worker uses the first selected card; several cards are not yet used together. The design for several cards gives each a share of every matrix's rows, which keeps each output element one card's exact dot product.
- No Vulkan, no GPU. Without a Vulkan loader, a usable card, or a passing verification, bankML runs on the CPU. Software renderers are never used.
- Remote cards are listed, never run.
bankml gpu --remoteneedscurland the network; nothing is provisioned. - No discrete card has been measured yet, and llama.cpp's own Vulkan build has not been measured as a comparison (TODO.md).
See also
- usage.md β
bankml gpu,BANKML_GPU,BANKML_GPU_SHARE,BANKML_GPU_LIMIT - sys.md and metrics.md β
/bankml/usage(GPU busy, VRAM, GTT, the limiter) and the answers' measurements - oracles.md β 0.2.13 and 0.2.14, the GPU oracles
- PERFORMANCE.md β decode speed
- huggingface.md β the rented GPUs
- TODO.md β 0.5.0, hardware
- forward.md β
Weights::openandWeights::mv - q1_0.md β the CPU kernel the GPU reproduces