CLM-v0.1-8B for vllm.cpp
This repository holds Contrastive-LM/CLM-v0.1-8B
converted into the single-directory layout that
vllm.cpp loads as ClmModel.
CLM is a bi-encoder System 1 decision model. A frozen Qwen3-8B backbone encodes
the state and each candidate answer separately. The last token's hidden state is
L2-normalized and projected by one of two MLP heads (a state head and an action
head) into a 512-dimensional space. A question's answer distribution is
softmax(scale * cos(state, candidate) / temperature). It answers typed
choice, noul and score questions and does not generate text.
This checkpoint only works with vllm.cpp (directly, or through LocalAI's
vllm-cpp backend). It is not a general-purpose checkpoint: transformers,
vLLM and llama.cpp do not know the ClmModel architecture or head.safetensors.
Redistribution and license
This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained or changed here.
- CLM heads: Contrastive-LM/CLM-v0.1-8B
@
e939398d4556fcd9400c76fa8c5a513202f42b0a(CLM_v0.1-8B.pt, sha256b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5), Apache-2.0. Reference code: Contrastive-LM/CLM @bb42c6c5bf914fd449bed2f6ca65be80602cb1f7. - Base backbone and tokenizer: Qwen/Qwen3-8B
@
b968826d9c46dd6066d109eabc6255188de91218, by the Qwen team, Apache-2.0. The base shards and tokenizer files are copied unchanged.
The upstream licenses apply to these files. Credit for the model belongs to the original authors.
Conversion
The converter is scripts/convert-clm.py from vllm.cpp (added in
a19294a9a). The CLM repository holds only the heads, as a torch pickle, and no
tokenizer, so the converter writes one directory: the base shards and tokenizer
files copied unchanged, head.safetensors with the checkpoint's own tensor names
(state_head.inp.weight, state_head.hidden.0.weight, state_head.norms.0.weight,
state_head.out.weight, their biases, and the same for action_head, all F32),
and a config.json that names ClmModel and carries the head config and the raw
logit_scale as clm_* keys. The pickle is loaded with weights_only=True.
hf download Contrastive-LM/CLM-v0.1-8B --revision e939398d4556fcd9400c76fa8c5a513202f42b0a \
--local-dir CLM-v0.1-8B
hf download Qwen/Qwen3-8B --revision b968826d9c46dd6066d109eabc6255188de91218 \
--local-dir Qwen3-8B
python3 scripts/convert-clm.py CLM-v0.1-8B --base-model-dir Qwen3-8B \
--output-dir clm-v0.1-8b
Head configuration: width 1536, depth 3, projection 512, GELU, LayerNorm, no
residual, logit_scale 4.6132 (scale min(exp(logit_scale), 100) = 100.0).
Files
| file | bytes | content |
|---|---|---|
model-0000N-of-00005.safetensors |
16,381,516,776 total | Qwen3-8B backbone, BF16, unchanged |
model.safetensors.index.json |
32,878 | shard index |
head.safetensors |
75,552,240 | state and action MLP heads, F32 |
config.json |
1,247 | Qwen3-8B config with architectures: ["ClmModel"] and the clm_* keys |
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt |
Qwen3-8B tokenizer, unchanged | |
generation_config.json |
239 | Qwen3-8B generation config, unchanged |
Serve it
With the vllm.cpp server:
build/examples/vllm-server --model clm-v0.1-8b --port 8000
LocalAI's vllm-cpp backend loads the same directory with
known_usecases: [decisions]. This checkpoint has not been run through LocalAI
yet; only the vllm.cpp server was used to verify it.
Then send a SystemOne request to POST /v1/systemone:
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "john works at google",
"questions": {"pick": {"type": "choice", "instructions": "entity type",
"criteria": {"person": null, "organization": null}}}}'
A response from this converted checkpoint, served on CPU by vllm.cpp:
{"model":"converted","answers":{"pick":{"type":"choice","choice":"person","confidence":0.9009225860482198,"probabilities":{"person":0.9504612930241099,"organization":0.04953870697589}}},"usage":{"billing_units":1,"input_tokens":9,"output_tokens":0},"latency_ms":2282.28}
The state head sees the state, a blank line, then the question's instructions.
Put the question in instructions. See the vllm.cpp recipe page docs/models/clm.md
for the exact candidate text each question type produces.
What was verified
On CPU, the converted checkpoint was served by vllm-server and compared with
the reference Engine.answer on 5 requests with 8 questions (choice, noul, score,
an object state, and a temperature override). The reference ran its own heads
over Qwen3-8B run by transformers in bf16.
| comparison | answers agree | max probability difference |
|---|---|---|
| this engine vs the reference (bf16) | 8 of 8 | 0.029 |
| this engine vs the reference (fp32) | 8 of 8 | 0.077 |
| the reference in fp32 vs the reference in bf16 | 8 of 8 | 0.049 |
The scale of 100 makes the answers sensitive to the encoder's rounding: the reference itself moves by up to 0.049 between bf16 and fp32.
What was not verified
- The reference's own path runs Qwen3-8B through vLLM's pooling server on a GPU.
The comparison above used
transformerson CPU in its place. - No CUDA run, no GGUF quantization, and no accuracy benchmark on any task.
- The reference's action-embedding cache is not implemented.
- The numbers above come from 8 questions. They show agreement with the reference, not accuracy of the model.
- Downloads last month
- -