CLM-v0.1-8B for vllm.cpp

This repository holds Contrastive-LM/CLM-v0.1-8B converted into the single-directory layout that vllm.cpp loads as ClmModel.

CLM is a bi-encoder System 1 decision model. A frozen Qwen3-8B backbone encodes the state and each candidate answer separately. The last token's hidden state is L2-normalized and projected by one of two MLP heads (a state head and an action head) into a 512-dimensional space. A question's answer distribution is softmax(scale * cos(state, candidate) / temperature). It answers typed choice, noul and score questions and does not generate text.

This checkpoint only works with vllm.cpp (directly, or through LocalAI's vllm-cpp backend). It is not a general-purpose checkpoint: transformers, vLLM and llama.cpp do not know the ClmModel architecture or head.safetensors.

Redistribution and license

This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained or changed here.

  • CLM heads: Contrastive-LM/CLM-v0.1-8B @ e939398d4556fcd9400c76fa8c5a513202f42b0a (CLM_v0.1-8B.pt, sha256 b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5), Apache-2.0. Reference code: Contrastive-LM/CLM @ bb42c6c5bf914fd449bed2f6ca65be80602cb1f7.
  • Base backbone and tokenizer: Qwen/Qwen3-8B @ b968826d9c46dd6066d109eabc6255188de91218, by the Qwen team, Apache-2.0. The base shards and tokenizer files are copied unchanged.

The upstream licenses apply to these files. Credit for the model belongs to the original authors.

Conversion

The converter is scripts/convert-clm.py from vllm.cpp (added in a19294a9a). The CLM repository holds only the heads, as a torch pickle, and no tokenizer, so the converter writes one directory: the base shards and tokenizer files copied unchanged, head.safetensors with the checkpoint's own tensor names (state_head.inp.weight, state_head.hidden.0.weight, state_head.norms.0.weight, state_head.out.weight, their biases, and the same for action_head, all F32), and a config.json that names ClmModel and carries the head config and the raw logit_scale as clm_* keys. The pickle is loaded with weights_only=True.

hf download Contrastive-LM/CLM-v0.1-8B --revision e939398d4556fcd9400c76fa8c5a513202f42b0a \
  --local-dir CLM-v0.1-8B
hf download Qwen/Qwen3-8B --revision b968826d9c46dd6066d109eabc6255188de91218 \
  --local-dir Qwen3-8B
python3 scripts/convert-clm.py CLM-v0.1-8B --base-model-dir Qwen3-8B \
  --output-dir clm-v0.1-8b

Head configuration: width 1536, depth 3, projection 512, GELU, LayerNorm, no residual, logit_scale 4.6132 (scale min(exp(logit_scale), 100) = 100.0).

Files

file bytes content
model-0000N-of-00005.safetensors 16,381,516,776 total Qwen3-8B backbone, BF16, unchanged
model.safetensors.index.json 32,878 shard index
head.safetensors 75,552,240 state and action MLP heads, F32
config.json 1,247 Qwen3-8B config with architectures: ["ClmModel"] and the clm_* keys
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt Qwen3-8B tokenizer, unchanged
generation_config.json 239 Qwen3-8B generation config, unchanged

Serve it

With the vllm.cpp server:

build/examples/vllm-server --model clm-v0.1-8b --port 8000

LocalAI's vllm-cpp backend loads the same directory with known_usecases: [decisions]. This checkpoint has not been run through LocalAI yet; only the vllm.cpp server was used to verify it.

Then send a SystemOne request to POST /v1/systemone:

curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "john works at google",
  "questions": {"pick": {"type": "choice", "instructions": "entity type",
    "criteria": {"person": null, "organization": null}}}}'

A response from this converted checkpoint, served on CPU by vllm.cpp:

{"model":"converted","answers":{"pick":{"type":"choice","choice":"person","confidence":0.9009225860482198,"probabilities":{"person":0.9504612930241099,"organization":0.04953870697589}}},"usage":{"billing_units":1,"input_tokens":9,"output_tokens":0},"latency_ms":2282.28}

The state head sees the state, a blank line, then the question's instructions. Put the question in instructions. See the vllm.cpp recipe page docs/models/clm.md for the exact candidate text each question type produces.

What was verified

On CPU, the converted checkpoint was served by vllm-server and compared with the reference Engine.answer on 5 requests with 8 questions (choice, noul, score, an object state, and a temperature override). The reference ran its own heads over Qwen3-8B run by transformers in bf16.

comparison answers agree max probability difference
this engine vs the reference (bf16) 8 of 8 0.029
this engine vs the reference (fp32) 8 of 8 0.077
the reference in fp32 vs the reference in bf16 8 of 8 0.049

The scale of 100 makes the answers sensitive to the encoder's rounding: the reference itself moves by up to 0.049 between bf16 and fp32.

What was not verified

  • The reference's own path runs Qwen3-8B through vLLM's pooling server on a GPU. The comparison above used transformers on CPU in its place.
  • No CUDA run, no GGUF quantization, and no accuracy benchmark on any task.
  • The reference's action-embedding cache is not implemented.
  • The numbers above come from 8 questions. They show agreement with the reference, not accuracy of the model.
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mudler/CLM-v0.1-8B-vllm-cpp

Finetuned
Qwen/Qwen3-8B
Finetuned
(4)
this model