cmf-decision / GPU.md
infosave's picture
Cortiq 0.8.0: primary Metal/Vulkan charts, Cargo installation and GPU API guide
9987fca verified
|
Raw History Blame Contribute Delete
5.12 kB
# Metal and Vulkan: the same CMF, accelerated
The model file, skills and gates stay unchanged. The GPU prepares text features
and computes reconstruction errors; it does not generate an answer.
**Available in Cortiq 0.8.0 and newer.** Install or update from crates.io:
```bash
cargo install cortiq-cli --locked
```
Requires Rust 1.88 or newer. GPU support is included; CPU remains the runtime default.
![GPU latency, within-host comparisons; different request schedules and independent scales.](figures/gpu.svg)
## RTX PRO 4000 Blackwell Β· Vulkan
| Dataset | CPU p50, ms | GPU p50, ms | GPU p95, ms | GPU p99, ms |
|---|---:|---:|---:|---:|
| BANKING77 | 24.05–24.30 | 1.24–1.24 | 1.77–1.78 | 2.23–2.30 |
| CLINC150 | 21.07–21.11 | 1.18–1.20 | 1.41–1.43 | 1.56–1.59 |
| MASSIVE | 16.23–16.26 | 1.10–1.15 | 1.33–1.37 | 1.50–1.56 |
Two full runs, each with 3,080 BANKING77, 4,500 CLINC150 and 2,974 MASSIVE rows.
CPU baseline: native portable-release path on the same Xeon E5-2690 v4 server.
CPU references are computed first; after GPU warmup, GPU requests run serially
without a CPU reference calculation between requests. Ranges span both runs.
This is warm local latency, not HTTP, cold startup, concurrent throughput or a Jev test.
**Idle behavior matters:** a final-build diagnostic alternates CPU/GPU requests
on 400 deterministically sampled rows per dataset. The GPU waits while the CPU
computes its reference; these gaps allow the RTX to lower its memory clock.
| Dataset | Rows | CPU p50, ms | Vulkan p50 / p95, ms |
|---|---:|---:|---:|
| BANKING77 | 400 | 25.29 | 8.57 / 9.16 |
| CLINC150 | 400 | 21.50 | 5.02 / 8.84 |
| MASSIVE | 400 | 16.56 | 3.40 / 7.50 |
Do not substitute continuous-stream latency for sparse-traffic latency.
No clocks, power limits or background keep-alive loads were forced. These
diagnostic samples are separate from the full-corpus quality checks above.
## Apple M4 Β· Metal
Default alternating CPU/GPU protocol, two full runs:
| Dataset | CPU p50, ms | GPU p50, ms | GPU p95, ms | GPU p99, ms |
|---|---:|---:|---:|---:|
| BANKING77 | 3.45–3.46 | 2.57–2.58 | 3.75–4.43 | 6.02–7.54 |
| CLINC150 | 3.26–3.27 | 2.52–2.55 | 3.38–3.54 | 4.09–6.37 |
| MASSIVE | 2.66–2.68 | 2.28–2.31 | 3.07–3.86 | 3.37–7.17 |
Separate serial streams, also two full runs:
| Dataset | CPU p50, ms | GPU p50, ms | GPU p95, ms | GPU p99, ms |
|---|---:|---:|---:|---:|
| BANKING77 | 3.43–3.43 | 2.08–2.09 | 3.13–3.15 | 3.75–3.80 |
| CLINC150 | 3.24–3.25 | 2.06–2.13 | 2.93–4.11 | 3.15–5.33 |
| MASSIVE | 2.63–2.64 | 1.81–3.21 | 2.70–5.44 | 2.92–5.97 |
The Mac had an active desktop. Metal shares the GPU with other applications;
one MASSIVE serial-stream run was slower than CPU, and some p95/p99 tails also
regressed. Acceleration is not a guarantee for every request or workload.
CPU remains the default; GPU selection is explicit.
## Correctness and resources
Across the two full runs per mode, winners, abstentions and hash features match
CPU exactly. This preserves the existing model's quality; it is not a new accuracy
improvement. Embedding tolerance remains `1e-5`; scaled reconstruction-error
difference is limited to `2e-5`. No retraining, oracle calls or lower precision.
[All per-dataset measurements and numerical checks](evidence/gpu.json).
The working path is used by CLI, `/v1/decisions` and `/v1/route`. Hardware tests cover
empty/long/Unicode input, non-aligned dimensions, non-orthogonal bases, multiple
skills and concurrent requests. GPU failures return errors, not a hidden CPU fallback.
Resident GPU weights use additional memory; model file size is not RAM or VRAM.
No energy comparison or hardware-normalized Jev resource comparison was measured.
| Host | CPU RSS, MiB | GPU-mode RSS, MiB |
|---|---:|---:|
| Apple M4 | 478.3 | 808.1 |
| Xeon + RTX | 437.2 | 648.3 |
RSS snapshots of the downloaded Cortiq 0.8.0 release binaries after 11 successful
HTTP requests per backend, all three skills
loaded; not peak memory. The Vulkan process also reported **522 MiB of VRAM**.
Four concurrent clients exercised routing; both decision API aliases were
checked. Temporary servers were stopped. CPU/GPU timing excludes model loading,
shader compilation, response serialization, network and external oracles.
## Choose a GPU
```bash
# Apple Silicon
CORTIQ_DECISION_DEVICE=metal cortiq decide cortiq-decision.cmf \
--skill banking77 -p "I still have not received my new card"
# Linux with a hardware Vulkan driver
CORTIQ_DECISION_DEVICE=vulkan CORTIQ_DECISION_VULKAN_ADAPTER="RTX PRO 4000" \
cortiq decide cortiq-decision.cmf \
--skill banking77 -p "I still have not received my new card"
```
The same environment variables apply to `cortiq serve`. `/healthz` reports the
actual adapter and completed GPU submissions. `gpu` timing includes the joint
encoder and reconstruction pass plus its wait, not a fictional kernel-only time.
Other GPUs require their own validation; a software Vulkan adapter is rejected.
[Back to model card](README.md) Β· [Existing Jev comparison](BENCHMARKS.md)