Split K/V cache quant, 16K ctx, cache-reuse, generation smoke test 4548f4e verified Leon4gr45 commited on 28 days ago
safety: cap threads at CPU_THREADS_MAX (default 2) to prevent oversubscription if detection over-reports 911ccc5 verified Leon4gr45 commited on 29 days ago
feat: cgroup-aware vCPU auto-detection (respects container CPU limit; no CPU_THREADS env needed) eb5a9f7 verified Leon4gr45 commited on 29 days ago
port verified llama.cpp CPU proxy: LFM2.5-VL-1.6B vision via mmproj, native OpenAI multimodal API, prebuilt binary 4baaa2f verified Leon4gr45 commited on 29 days ago