File size: 21,071 Bytes
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
# `bankML/gpu/` β€” the video-card component: find GPUs, verify them, and give a verified card a share of the 1-bit matmuls

## Summary

`bankML/gpu/` finds every GPU on the machine, describes it, and puts a card to work only after it has proved that it
gives the CPU kernels' bits. It is a registry of backends. Each backend is a module with one discovery function, and
nothing is linked at build time: the Vulkan backend opens its driver library at run time with `dlopen`, so bankML builds
and runs everywhere, and a machine without the library simply has no devices from that backend.

A card must be a real GPU (integrated or discrete) with a compute queue. Software renderers such as Mesa's llvmpipe run
on the CPU and are refused. Before a card computes anything it must pass the on-card oracle, `kernels::verify_q1_0`
(the oracle rule of [TECHNICAL.md Β§III.4](../TECHNICAL.md): the same bits first, then speed).
Inside the forward pass a verified card takes the first, calibrated share of a 1-bit (`Q1_0`) matrix's rows while
the CPU pool computes the rest β€” for each matrix shape where it was measured to pay, within the limit
`BANKML_GPU_LIMIT` sets (0.3.7). Each output element stays one device's exact dot product, so a GPU changes the speed,
never the tokens.

The component is reached three ways: the CLI (`bankml gpu`, `bankml gpu --remote`, `bankml gpu --verify`), the forward
pass (`Weights::open` opens a `gpu::worker::Worker` for `Q1_0` models), and the tests
(`gpu_q1_0_mat_vec_bit_exact`). Remote cards, the GPUs Hugging Face rents, are a separate kind: listed on request,
never selected or started.

## Technical usage

### CLI

```sh
bankml gpu              # every video card found (Vulkan, merged with /sys/class/drm) and which bankML will use, as JSON
bankml gpu --remote     # the same, plus the GPUs Hugging Face rents (listed only)
bankml gpu --verify     # the bit-exact kernel oracle on each selected card; exit code 2 if any card fails
```

`bankml gpu` prints `report_json(remote)`: a `devices` array, the `selected` cards, the `notes` (why a backend found
nothing), and a `kernels` field (Q1_0 on a verified card; Q2_0 and F16 on the CPU). With `--verify`, each selected card prints one line per kernel and shape, or
`NOT VERIFIED β€” <reason>; bankml will not use this card`. With no usable card it prints
`bankml gpu --verify: no usable GPU found; bankml runs on the CPU`.

### Environment

| variable | effect |
|---|---|
| `BANKML_GPU=off` (also `none`, `cpu`) | turns the component off: nothing is selected, the worker is not opened |
| `BANKML_GPU=0,2` | picks backend indices (the `index` that `bankml gpu` reports) |
| `BANKML_GPU_SHARE=0.3` | overrides the calibrated share of each 1-bit matrix's rows; clamped to 0–0.95 |
| `BANKML_GPU_LIMIT=0.8` | the share of the card's memory and time bankML may use (0.3.7; clamped 0.05–1, default 0.8); see below |

### `mod.rs` β€” the registry

```rust
pub type Discover = fn() -> Result<Vec<Device>, String>;
pub const BACKENDS: &[(&str, Discover)] = &[("vulkan", vulkan::devices)];
pub const REMOTE_BACKENDS: &[(&str, Discover)] = &[("huggingface", hf::devices)];

pub fn discover() -> (Vec<Device>, Vec<String>)
pub fn selected(devs: &[Device]) -> Vec<Device>
pub fn report_json(remote: bool) -> String
```

- `Device` holds the backend's view (name, vendor and device ids, `Kind`, API version, `device_local_bytes`,
  `host_visible_device_local`, `compute_queues`) plus `count` and `usd_per_hour` for rented hardware, and an optional
  `Sysfs` record.
- `Kind` is `Integrated`, `Discrete`, `Virtual`, `Cpu` (a software renderer, never used), `Remote` (rented, never
  started automatically) or `Other`.
- `Device::usable()` is true only for `Integrated` or `Discrete` with at least one compute queue.
- `discover()` asks every local backend and attaches the kernel's view of the card from `/sys/class/drm/card*/device`:
  driver, VRAM and GTT totals, PCI address, NUMA node. A card is matched by vendor and device id, so it is described
  by what the driver says, not only by what the API reports.
- `device_local_bytes` is the sum of the device-local heaps, the memory the device addresses fastest.
  `host_visible_device_local` is true when some device-local memory is also host-visible (integrated memory,
  resizable BAR), so weights need no staging copy.
- `selected()` applies `BANKML_GPU`, keeps the usable devices, and orders them discrete first, then by device-local
  memory, largest first.
- A new backend (CUDA, ROCm, Metal, …) is a new module and one line in `BACKENDS`.
- A remote card is never selected automatically, because it costs money.

### `vulkan.rs` β€” Vulkan through `dlopen`

```rust
pub fn devices() -> Result<Vec<Device>, String>
pub(super) fn instance() -> Result<(GetInstanceProcAddr, *mut c_void), String>
```

One API covers every GPU a Vulkan driver exposes: AMD (RADV, AMDVLK), NVIDIA, Intel (ANV), Arm, Qualcomm and Apple
(MoltenVK). The loader is opened at run time: `libvulkan.so.1`, `libvulkan.so`, `libvulkan.1.dylib`, then
`libMoltenVK.dylib`.
Every entry point comes from `vkGetInstanceProcAddr`. The few C structs bankML needs are declared in the file from the
Vulkan 1.x headers. `devices()` creates an instance, enumerates the physical devices, reads their properties, memory
heaps and queue families, and destroys the instance. With no loader it returns
`no Vulkan loader (libvulkan.so.1) on this machine`, and bankML carries on with the CPU. `instance()` creates the
instance compute uses and never destroys it; it lives for the rest of the process. The device name is read from the
first 276 bytes of `VkPhysicalDeviceProperties`, written into a 4 KiB buffer larger than the whole struct.

### `spirv.rs` β€” bankML's own SPIR-V assembler

bankML writes its compute shaders as SPIR-V words itself, so no shader compiler is needed at build or run time (the
development machine has none, and bankML takes no crates).
`Module` keeps the sections apart and joins them in the order the spec requires (`Module::words()`). It covers only
what compute kernels use: scalar and vector types, storage buffers, push constants, the invocation ids, workgroup
memory and a control barrier, integer and float arithmetic, GLSL.std.450 `Fma`, and structured control flow
(selections and loops). Steps the CPU kernels fuse are written as explicit `Fma`.

```rust
pub fn op(&mut self, opcode: u16, result_ty: u32, operands: &[u32]) -> u32
pub fn two_sum(&mut self, f32t: u32, a: u32, b: u32) -> (u32, u32)
pub fn fma_exact(&mut self, tys: (u32, u32, u32), a: u32, b: u32, c: u32, c_mask: u32, c1: u32, c0: u32, cf0: u32) -> u32
```

- `op` decorates every `OpFMul`, `OpFAdd`, `OpFSub` and extended instruction `NoContraction`, so a driver cannot fuse
  what the CPU kernels keep separate.
- `fma_exact` builds a correctly rounded `fma(a, b, c)` from plain multiplies and adds, whatever the driver does with
  `Fma`. `a` must carry at most 12 significant bits (an f16 value does). `b` is split into two 12-bit halves so both
  partial products are exact, Knuth's TwoSum keeps every rounding error, and Boldo and Melquiond's
  `RN(th + RO(tl + ul))` (rounding to odd, 2008) gives the FMA's single rounding.
- Why it exists: RADV on the Radeon Vega 3 does not fuse `Fma`. On real layer weights the 0.2.13 kernels differed
  from the CPU by about one ulp on 70 % of rows, and decorating the `Fma` `NoContraction` did not change it
  (CHANGELOG 0.2.14).

### `kernels.rs` β€” the Q1_0 kernels

```rust
pub const Q1_0_BINDINGS: u32 = 5;
pub const LOCAL_SIZE: u32 = 64;
pub fn q1_0_mat_vec() -> Vec<u32>
pub fn q1_0_mat_vec8() -> Vec<u32>
pub fn verify_q1_0(gpu: &super::compute::Gpu) -> Result<Vec<String>, String>
pub fn pack_q1_0(w: &[u8], rows: usize, n: usize) -> (Vec<f32>, Vec<u32>)
pub fn pack_act(a: &crate::q1_0::Q8Act) -> (Vec<f32>, Vec<i32>)
```

- Both kernels follow `q1_0::vec_dot_ref` step by step: per 128-weight block, four 32-element sub-blocks; eight lanes
  of exact integer sums (ggml's i8 negation wraps βˆ’128 to βˆ’128); `ab = d1Β·s` on the first sub-block and
  `fma(d1, s, ab)` after; the outer `fma(d0, ab, acc)` through `fma_exact`; then the horizontal sum
  `(a0+a4 + a2+a6) + (a1+a5 + a3+a7)`.
- `q1_0_mat_vec` uses one invocation per output row (dispatch `ceil(rows / 64)` workgroups). `q1_0_mat_vec8` uses
  eight per row, one per accumulation lane, which is valid because each lane's chain over the blocks is independent
  in the CPU kernel too; a workgroup of 64 covers 8 rows. The lane sums meet in workgroup memory and lane 0 adds them
  in the CPU's order, giving the same bits as `q1_0_mat_vec` with eight times the parallelism. Dispatch
  `ceil(rows Β· 8 / 64)` workgroups.
- Bindings: weight scales (f32), weight bits (u32, four per block), activation scales (f32 per 32), activation quants
  (i32 words of four i8), output (f32). Push constants: rows, blocks per row.
- `pack_q1_0` repacks weights once (scales f16β†’f32, exact; bits as u32 words). `pack_act` repacks the q8_0 activation
  per call. The numbers are unchanged, only aligned.

### `compute.rs` β€” Vulkan compute

```rust
impl Gpu {
    pub fn open(index: usize) -> Result<Gpu, String>
    pub fn buffer(&self, bytes: usize) -> Result<Buffer, String>
    pub fn upload<T: Copy>(&self, data: &[T]) -> Result<Buffer, String>
    pub fn write<T: Copy>(&self, b: &Buffer, data: &[T])
    pub fn read_f32(&self, b: &Buffer, n: usize) -> Vec<f32>
    pub fn allocated(&self) -> u64                    // bytes bankML's buffers hold (0.3.7, for the limiter)
    pub fn free(&self, b: Buffer)
    pub fn pipeline(&self, spirv: &[u32], bindings: u32, push_bytes: u32) -> Result<Pipeline, String>
    pub fn run(&self, p: &Pipeline, bufs: &[&Buffer], push: &[u32], groups: u32) -> Result<(), String>
    pub fn submit(&self, p: &Pipeline, bufs: &[&Buffer], push: &[u32], groups: u32) -> Result<(), String>
    pub fn wait(&self) -> Result<(), String>
}
```

A device with one compute queue, host-visible coherent buffers mapped for their whole life (device-local and
host-visible memory when the card has it, as an integrated GPU does), pipelines from bankML's own SPIR-V, and a
dispatch that waits on a fence. `run` is `submit` then `wait`. One submission is in flight at a time (one command
buffer, one fence, one descriptor set per pipeline), so `wait` must follow each `submit` before the next; `wait`
without a pending `submit` blocks indefinitely. Every Vulkan call's result is checked.

On a card without device-local host-visible memory, buffers live in host-visible memory the CPU can write. `Gpu` is
`Send` but not `Sync`: one thread drives it at a time (the forward pass holds the worker behind a `Mutex`). Nothing is
destroyed on drop: a `Buffer` is released only by `Gpu::free`, pipelines are never destroyed, and dropping a `Gpu`
waits for the device to go idle and leaves the device objects for the driver to reclaim at process exit.

### `worker.rs` β€” the GPU beside the CPU's threads

```rust
impl Worker {
    pub fn open(pool: &Pool) -> Result<Option<Worker>, String>
    pub fn begin(&mut self, name: &str, w: &[u8], rows: usize, a: &Q8Act) -> Result<usize, String>
    pub fn finish(&mut self, out: &mut [f32]) -> Result<(), String>
}
```

- `open` takes the first selected card, runs `verify_q1_0` on it, and builds the `q1_0_mat_vec8` pipeline. It
  returns `None` when there is no usable card or the share is below 0.02.
- The share comes from a calibration: the same 4096Γ—4096 product timed on the card and on the CPU pool (best of
  three each), `share = card rate / (card rate + CPU rate)`. `BANKML_GPU_SHARE` overrides it.
- In `Weights::mv`, for each `Q1_0` product, `begin` starts the card on the first `share` of the rows and returns
  how many it took; the CPU pool computes the rest with `q1_0::mat_vec_par`; `finish` waits and copies the card's
  rows in. `finish` is called after every `begin`, even when the card took no rows, because it also records the
  shape's timing.
- Only the card's share of each matrix is copied to it, repacked exactly, on first use. Activations wider than
  16,384 are left to the CPU.
- A failure to open or verify is logged (`bankml: GPU not used: …`) and the CPU path runs alone.

### `hf.rs` β€” Hugging Face's rented GPUs

```rust
pub const HARDWARE_URL: &str = "https://huggingface.co/api/jobs/hardware";
pub fn devices() -> Result<Vec<Device>, String>
pub fn parse(text: &str) -> Result<Vec<Device>, String>
```

Reads the public Jobs hardware list (no token) through the system `curl`, because bankML has no TLS stack of its own.
Each GPU flavor becomes one `Kind::Remote` device with its card count, memory and price per hour (the API's
per-minute price Γ— 60). A multi-card flavor such as `h200x8` is one device with `count` 8. Nothing is started, rented
or paid for. Using a rented card means running bankML there as a Hugging Face Job, where the Vulkan backend finds it
like any local card. See [huggingface.md](../huggingface.md).

The flavors range from an Nvidia T4 to 8Γ— H200. Hugging Face provisions them on the large clouds (its Inference
Endpoints name AWS, Azure and GCP); NVIDIA agreed on 2026-09-02 to acquire Hugging Face. `curl` is used as the
installer's downloads use it; without `curl` or the network the backend reports why and bankML carries on with local
cards. A price per hour is reported only when the API's price is per minute.

### The limiter and per-shape calibration (0.3.7)

```rust
pub fn limit() -> f64                 // BANKML_GPU_LIMIT, clamped 0.05–1, default 0.8
pub fn status() -> Option<Status>     // None when no card works in this process
pub fn status_json() -> String        // GET /bankml/usage's "gpu_limiter"
```

- **`BANKML_GPU_LIMIT`** = L: bankML's buffers stay within L of the heap they come from (the `Gpu.heap_bytes` field,
  against `Gpu::allocated()`), and on an integrated card within L of the RAM they could use (what they hold plus what
  the system has available). After a dispatch of `d` the card rests `dΒ·(1βˆ’L)/L`. A matrix that does not fit, or a
  product arriving while the card rests, runs whole on the CPU β€” never a wait, the same bits.
- **Per-shape calibration**: each matrix shape's first products alternate between the calibrated share and the CPU
  alone, timed `begin` β†’ `finish` in the real pipeline (the CPU's rows included). After a warm-up (which includes
  the upload) and 6 timings each (`TRIALS`) the faster median is kept, and a shape decided for the CPU frees its
  buffers. A product the card skipped because it was resting or out of memory is not counted as a timing of the
  card. The bits are the same either way.
- Why per shape: the one-off calibration measures a 4096Γ—4096 product in isolation. In the forward pass a smaller
  matrix may not repay the submit-and-wait, and on an integrated card (whose heap is system RAM) the card and the CPU
  share the memory bandwidth that decode is bound by.
- `status_json` reports `card`, `limit`, `share`, `allocated_bytes`, `heap_bytes`, `integrated`, `busy`,
  `products_on_card`, `products_on_cpu_resting`, `products_on_cpu_memory`, `shapes_on_card`, `shapes_on_cpu` and
  `shapes_tuning`. The console's Admin tab charts it ([console.md](console.md)).
- Measured (CHANGELOG 0.3.7; Vega 3, Bonsai-1.7B, interleaved A/B, 4 rounds, the same answer every time): with one
  global share the card cost decode 18 % (10.40 β†’ 8.51–8.58 tok/s); per shape, 10.36 tok/s off against 10.30–10.38
  on β€” neutral β€” with ~13 MB on the card instead of 66 MB.

## How it is verified

- **`gpu_q1_0_mat_vec_bit_exact`** (`kernels.rs`, ignored by default: needs a Vulkan GPU; in the release gate). Both
  kernels on every usable card against bankML's CPU kernel, which is bit-exact against ggml, on random weights and
  activations with βˆ’128 quants. Shapes 64Γ—128 to 12288Γ—4096. On the Vega 3, every row of every shape is bit-exact
  ([oracles.md](../oracles.md), 0.2.13).
- **`bankml gpu --verify`** runs `verify_q1_0`: the same comparison, plus a layer-shaped regime (scales and magnitudes
  that vary per block, so products are inexact in f32). With the driver's `Fma` put back, it refuses the Vega 3
  ("row 1 differs … bankml will not use this card"). The worker runs the same check before it takes any work.
- **The forward pass with the card working.** Every token oracle runs through `Weights::mv`, so with a verified card
  present they all run with the GPU computing its share. In 0.2.14 the whole 1-bit model passed at 1,064 of 1,064
  rows with the Vega 3 on 26 % of every matrix, and the greedy and sampling checks went through the same path.
- Unit tests: `software_renderers_are_never_selected_and_discrete_cards_come_first`, `discovery_never_panics`
  (`mod.rs`), `a_multi_card_flavor_is_one_device_with_its_count` (`hf.rs`), `header_and_string_words` (`spirv.rs`).

## Advantages and efficiency

- **The same bits, then speed.** A card is trusted only after it reproduces the CPU kernel's bits on the card, and it
  works on whole rows, so no element is split between devices. The answer and the receipt do not depend on whether a
  GPU was present.
- **No SDK, no crate, no shader compiler.** Vulkan is opened with `dlopen` and every entry point comes from
  `vkGetInstanceProcAddr`; shaders are assembled in Rust by `spirv.rs`. The same binary runs on a machine with or
  without a GPU. One API covers AMD, NVIDIA, Intel, Arm and Qualcomm, and a rented NVIDIA card needs no CUDA toolkit
  ([huggingface.md](../huggingface.md)).
- **The card and the CPU finish together.** The calibrated share gives the card the fraction of rows that matches its
  rate, and the CPU pool computes the rest while the card works (`submit` returns at once; `wait` is called after the
  CPU part).
- **Exact FMA only where needed.** The inner `fma(d1, s, ab)` stays a plain `Fma`, because `d1Β·s` is exact in f32;
  only the outer `fma(d0, ab, acc)` uses `fma_exact` (CHANGELOG 0.2.14).
- **Weights copied once, and only the card's share.** On an integrated GPU, buffers are allocated in device-local,
  host-visible memory when the card offers it.
- **Measured.** On the Vega 3 the eight-lane kernel matches one CPU thread, 1.7–2.0 ms per 4096Γ—4096 (CHANGELOG
  0.2.13). The calibrated share is 26–35 % there. Decode is unchanged within noise: 1.93–2.00 tokens/s with the card
  against 1.96–1.97 without (CHANGELOG 0.2.14).
- **Rust practice visible in the code.** No external crate (the manifest's `[dependencies]` is empty). `unsafe` is
  confined to the FFI files `vulkan.rs` and `compute.rs`, with `SAFETY` comments, behind safe methods on `Gpu`; the
  registry, kernels, assembler, worker and Hugging Face backend have none. Errors are `Result<_, String>` with the
  failing call named, and a failure falls back to the CPU instead of stopping. Verification fails closed. The
  toolchain is pinned to Rust 1.99.0 in `rust-toolchain.toml`.
- **Where optimization goes next** ([TODO.md](../TODO.md), 0.5.0): batched submissions (Q/K/V and gate/up in one
  command buffer, one wait per group, persistent descriptor sets), the Q2_0 (ternary) GPU kernel in `--verify`,
  several cards each taking a share of every matrix's rows, and a discrete-card measurement (a rented T4 or L4 run,
  only with the owner's approval).

## Limitations

- **Speed-neutral on the APU.** On the Vega 3 the card is worth about one CPU core, and each of the 253 matrix calls
  per token pays one submit and wait, which cancels the gain (CHANGELOG 0.2.14). Per-shape calibration (0.3.7) keeps
  the card from costing decode, but does not make it faster there. One submission is in flight at a time.
- **`Q1_0` only.** The worker is opened only for 1-bit models; there is no Q2_0 (ternary) or F16 GPU kernel yet.
- **One card.** The worker uses the first selected card; several cards are not yet used together. The design for
  several cards gives each a share of every matrix's rows, which keeps each output element one card's exact dot
  product.
- **No Vulkan, no GPU.** Without a Vulkan loader, a usable card, or a passing verification, bankML runs on the CPU.
  Software renderers are never used.
- **Remote cards are listed, never run.** `bankml gpu --remote` needs `curl` and the network; nothing is provisioned.
- No discrete card has been measured yet, and llama.cpp's own Vulkan build has not been measured as a comparison
  ([TODO.md](../TODO.md)).

## See also

- [usage.md](../usage.md) β€” `bankml gpu`, `BANKML_GPU`, `BANKML_GPU_SHARE`, `BANKML_GPU_LIMIT`
- [sys.md](sys.md) and [metrics.md](metrics.md) β€” `/bankml/usage` (GPU busy, VRAM, GTT, the limiter) and the answers' measurements
- [oracles.md](../oracles.md) β€” 0.2.13 and 0.2.14, the GPU oracles
- [PERFORMANCE.md](../PERFORMANCE.md) β€” decode speed
- [huggingface.md](../huggingface.md) β€” the rented GPUs
- [TODO.md](../TODO.md) β€” 0.5.0, hardware
- [forward.md](forward.md) β€” `Weights::open` and `Weights::mv`
- [q1_0.md](q1_0.md) β€” the CPU kernel the GPU reproduces