kev-0.8b for vllm.cpp
This repository holds jaredpalmer/kev-0.8b
converted into the single-directory layout that
vllm.cpp loads as KevModel.
kev is a System 1 decision model. It answers typed choice, noul and
score questions about a text state with one scoring pass per question. It
does not generate text. kev is a frozen Qwen/Qwen3.5-0.8B-Base backbone, a
rank-16 LoRA adapter, and a PointerHead readout.
This checkpoint only works with vllm.cpp (directly, or through LocalAI's
vllm-cpp backend). It is not a general-purpose checkpoint: transformers,
vLLM and llama.cpp do not know the KevModel architecture or head.safetensors.
Redistribution and license
This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained or changed here, other than the LoRA merge described below.
- kev adapter and PointerHead: jaredpalmer/kev-0.8b
@
9a45d25eb2ab761841196625383fa1dff0e56c1e, by Jared Palmer, Apache-2.0. Reference code: jaredpalmer/kev. - Base backbone and tokenizer: Qwen/Qwen3.5-0.8B-Base
@
dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68, by the Qwen team, Apache-2.0.
The upstream licenses apply to these files. Credit for the model belongs to the original authors.
Conversion
The converter is scripts/convert-kev.py from vllm.cpp at commit
96788348627b6a079fcc3ef7fc6676a970965d7b. It merges the LoRA once
(W' = W + 2.0 * (B @ A), r=16, alpha=32, math in F32, stored in the base
dtype BF16), drops the base's vision and MTP tensors, converts head.pt to
head.safetensors, and writes a config.json with architectures: ["KevModel"].
hf download jaredpalmer/kev-0.8b --revision 9a45d25eb2ab761841196625383fa1dff0e56c1e \
--local-dir kev-0.8b
hf download Qwen/Qwen3.5-0.8B-Base --revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68 \
--local-dir Qwen3.5-0.8B-Base
python3 scripts/convert-kev.py kev-0.8b \
--base-model-dir Qwen3.5-0.8B-Base --output-dir kev-0.8b-vllm-cpp
The converter logged: 186 of 186 LoRA modules merged, 168 vision/MTP tensors dropped, 320 tensors written. Environment: torch 2.11.0, safetensors 0.7.0, CPU.
Files
| file | bytes | content |
|---|---|---|
model.safetensors |
1,504,832,320 | merged Qwen3.5-0.8B backbone, BF16 |
head.safetensors |
2,099,504 | PointerHead q/k weight and bias, F32 |
config.json |
2,526 | base config, architectures: ["KevModel"], kev_head_dim: 256, kev_temperature: 2.3510958125672174 |
meta.json |
258 | head metadata from head.pt |
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt |
copied from the base model |
Serve it with vllm.cpp
Build the vllm.cpp server (-DVLLM_CPP_SERVER=ON, target vllm-server), then:
hf download mudler/kev-0.8b-vllm-cpp --revision c17e73666ded1e9d284470eae7e0de9a27294e77 --local-dir kev-0.8b-vllm-cpp
build/examples/vllm-server --model kev-0.8b-vllm-cpp --served-model-name kev --port 8000
The server logs kev decision model (KevModel); serving /v1/systemone.
Serve it with LocalAI
A model config for LocalAI's vllm-cpp backend:
name: kev-0.8b
backend: vllm-cpp
known_usecases:
- decisions
parameters:
model: mudler/kev-0.8b-vllm-cpp
artifacts:
- name: model
target: model
source:
type: huggingface
repo: mudler/kev-0.8b-vllm-cpp
revision: c17e73666ded1e9d284470eae7e0de9a27294e77
Then send the same request body to LocalAI's POST /v1/systemone with
"model": "kev-0.8b".
Example
Request (served on CPU from this exact directory, 2026-09-30):
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "kev",
"state": "I was charged twice for my subscription this month. Please refund the duplicate payment.",
"questions": {
"refund": {"type": "noul", "instructions": "Does the user request a refund?"},
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "Payments and refunds", "technical": "Software bugs", "sales": "New purchases"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Routine", "Urgent", "Emergency"]}}}'
Response (the rl_agent fields are omitted here):
{
"answers": {
"department": {"type": "choice", "choice": "billing", "confidence": 0.9264,
"probabilities": {"billing": 0.951, "sales": 0.0264, "technical": 0.0227}},
"refund": {"type": "noul", "noul": 0.9667},
"urgency": {"type": "score", "score": 0.7069, "confidence": 0.7244,
"legend": {"0": "Routine", "1": "Urgent", "2": "Emergency"},
"probabilities": {"0": 0.4222, "1": 0.4487, "2": 0.1291}}
},
"model": "kev",
"usage": {"input_tokens": 111, "output_tokens": 0},
"latency_ms": 2451.45
}
kev's answer semantics follow the kev reference: choice confidence is
(max(p) - 1/K) / (1 - 1/K), score is the expected level index with
confidence 1 - E|level - mode| / (L - 1), noul is p(true), and
probabilities are rounded to 4 decimals in this response.
What was verified and what was not
Verified for this upload:
- The conversion above ran at the stated vllm.cpp commit and merged all 186 LoRA modules.
- The directory loads in the vllm.cpp CPU server built from the same commit, and the one request above returns the answers shown. The answers are plausible for the input. This is a smoke test, not an accuracy measurement.
Recorded by the vllm.cpp project, not re-measured for this upload:
- PointerHead golden-vector tests (25 cases) and LoRA-merge / hidden-state
goldens against the kev Python reference at
jaredpalmer/kev@19dcae9b6e3e1a48200c5825aad9fc200d31e20a. - A 5-case end-to-end comparison (choice 2 and 3 options, score 5 levels, noul, multi-question) against the kev reference server, recorded as equal.
Not verified:
- No GPU (CUDA, ROCm, Metal) run of this checkpoint.
- No comparison of this upload against the kev reference server. No vLLM gate exists for kev, because vLLM does not serve this architecture.
- No accuracy benchmark.
- No GGUF or quantized variant.
- Downloads last month
- -
Model tree for mudler/kev-0.8b-vllm-cpp
Base model
Qwen/Qwen3.5-0.8B-Base