kev-0.8b for vllm.cpp

This repository holds jaredpalmer/kev-0.8b converted into the single-directory layout that vllm.cpp loads as KevModel.

kev is a System 1 decision model. It answers typed choice, noul and score questions about a text state with one scoring pass per question. It does not generate text. kev is a frozen Qwen/Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter, and a PointerHead readout.

This checkpoint only works with vllm.cpp (directly, or through LocalAI's vllm-cpp backend). It is not a general-purpose checkpoint: transformers, vLLM and llama.cpp do not know the KevModel architecture or head.safetensors.

Redistribution and license

This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained or changed here, other than the LoRA merge described below.

The upstream licenses apply to these files. Credit for the model belongs to the original authors.

Conversion

The converter is scripts/convert-kev.py from vllm.cpp at commit 96788348627b6a079fcc3ef7fc6676a970965d7b. It merges the LoRA once (W' = W + 2.0 * (B @ A), r=16, alpha=32, math in F32, stored in the base dtype BF16), drops the base's vision and MTP tensors, converts head.pt to head.safetensors, and writes a config.json with architectures: ["KevModel"].

hf download jaredpalmer/kev-0.8b --revision 9a45d25eb2ab761841196625383fa1dff0e56c1e \
  --local-dir kev-0.8b
hf download Qwen/Qwen3.5-0.8B-Base --revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68 \
  --local-dir Qwen3.5-0.8B-Base
python3 scripts/convert-kev.py kev-0.8b \
  --base-model-dir Qwen3.5-0.8B-Base --output-dir kev-0.8b-vllm-cpp

The converter logged: 186 of 186 LoRA modules merged, 168 vision/MTP tensors dropped, 320 tensors written. Environment: torch 2.11.0, safetensors 0.7.0, CPU.

Files

file bytes content
model.safetensors 1,504,832,320 merged Qwen3.5-0.8B backbone, BF16
head.safetensors 2,099,504 PointerHead q/k weight and bias, F32
config.json 2,526 base config, architectures: ["KevModel"], kev_head_dim: 256, kev_temperature: 2.3510958125672174
meta.json 258 head metadata from head.pt
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt copied from the base model

Serve it with vllm.cpp

Build the vllm.cpp server (-DVLLM_CPP_SERVER=ON, target vllm-server), then:

hf download mudler/kev-0.8b-vllm-cpp --revision c17e73666ded1e9d284470eae7e0de9a27294e77 --local-dir kev-0.8b-vllm-cpp
build/examples/vllm-server --model kev-0.8b-vllm-cpp --served-model-name kev --port 8000

The server logs kev decision model (KevModel); serving /v1/systemone.

Serve it with LocalAI

A model config for LocalAI's vllm-cpp backend:

name: kev-0.8b
backend: vllm-cpp
known_usecases:
  - decisions
parameters:
  model: mudler/kev-0.8b-vllm-cpp
artifacts:
  - name: model
    target: model
    source:
      type: huggingface
      repo: mudler/kev-0.8b-vllm-cpp
      revision: c17e73666ded1e9d284470eae7e0de9a27294e77

Then send the same request body to LocalAI's POST /v1/systemone with "model": "kev-0.8b".

Example

Request (served on CPU from this exact directory, 2026-09-30):

curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "kev",
  "state": "I was charged twice for my subscription this month. Please refund the duplicate payment.",
  "questions": {
    "refund": {"type": "noul", "instructions": "Does the user request a refund?"},
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "Payments and refunds", "technical": "Software bugs", "sales": "New purchases"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Routine", "Urgent", "Emergency"]}}}'

Response (the rl_agent fields are omitted here):

{
  "answers": {
    "department": {"type": "choice", "choice": "billing", "confidence": 0.9264,
                   "probabilities": {"billing": 0.951, "sales": 0.0264, "technical": 0.0227}},
    "refund": {"type": "noul", "noul": 0.9667},
    "urgency": {"type": "score", "score": 0.7069, "confidence": 0.7244,
                "legend": {"0": "Routine", "1": "Urgent", "2": "Emergency"},
                "probabilities": {"0": 0.4222, "1": 0.4487, "2": 0.1291}}
  },
  "model": "kev",
  "usage": {"input_tokens": 111, "output_tokens": 0},
  "latency_ms": 2451.45
}

kev's answer semantics follow the kev reference: choice confidence is (max(p) - 1/K) / (1 - 1/K), score is the expected level index with confidence 1 - E|level - mode| / (L - 1), noul is p(true), and probabilities are rounded to 4 decimals in this response.

What was verified and what was not

Verified for this upload:

  • The conversion above ran at the stated vllm.cpp commit and merged all 186 LoRA modules.
  • The directory loads in the vllm.cpp CPU server built from the same commit, and the one request above returns the answers shown. The answers are plausible for the input. This is a smoke test, not an accuracy measurement.

Recorded by the vllm.cpp project, not re-measured for this upload:

  • PointerHead golden-vector tests (25 cases) and LoRA-merge / hidden-state goldens against the kev Python reference at jaredpalmer/kev@19dcae9b6e3e1a48200c5825aad9fc200d31e20a.
  • A 5-case end-to-end comparison (choice 2 and 3 options, score 5 levels, noul, multi-question) against the kev reference server, recorded as equal.

Not verified:

  • No GPU (CUDA, ROCm, Metal) run of this checkpoint.
  • No comparison of this upload against the kev reference server. No vLLM gate exists for kev, because vLLM does not serve this architecture.
  • No accuracy benchmark.
  • No GGUF or quantized variant.
Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mudler/kev-0.8b-vllm-cpp

Finetuned
(122)
this model