OpenJev Flash 9B, FP8 checkpoint

The FP8 checkpoint of OpenJev Flash 9B: the same weights stored as FP8 (e4m3, 128×128 weight blocks, dynamic activations), 11.9 GB instead of 18.8 GB.

On par with Cloudflare's Clef-Flash and ahead of Kev-9B and Nimble 9B on JevBench. The results, the API and the use cases are on the main card.

  • Drop-in for vLLM: the main card's command and helper, pointed at this repository. The weights are already FP8, so nothing is quantized at load time.
  • Matches the 16-bit model: 80.0% on a fixed 1,789-question check, the same as the 16-bit weights, and 79.2% on the frozen 10,000 text questions against 79.4% for the 16-bit weights served in FP8.
  • 37% smaller download: 11.9 GB instead of 18.8 GB.

Accuracy and download size of the FP8 checkpoint against the 16-bit model

Measured

One GPU, vLLM 0.29. The 10,000 questions go through the OpenJev helper with the card's calibration; the 1,789-question check reads the letter probabilities directly.

test 16-bit weights 16-bit, served in FP8 this FP8 file
10,000 text questions, 34 public sources — 79.4% 79.2%
fixed 1,789-question check 80.0% 79.7% 80.0%

Run it

pip install "vllm==0.29.0" "openai==3.26.0" "httpx==0.28.1"
hf download openjev/OpenJev-Flash-9B helper/shim.py --local-dir openjev-flash-9b

Terminal 1, the model (downloads the weights on first start):

vllm serve openjev/OpenJev-Flash-9B-FP8 --host 127.0.0.1 --served-model-name qwen --port 8000 \
  --enable-prefix-caching --max-model-len 16384 --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image":1}' --trust-remote-code --max-num-seqs 16 \
  --max-logprobs 64 --gdn-prefill-backend triton --quantization fp8

Terminal 2, the decision API in front of it:

VLLM=http://127.0.0.1:8000/v1 TOKENIZER=openjev/OpenJev-Flash-9B-FP8 \
READOUT_T=1.07 READOUT_NOUL_T=1.074766 READOUT_NOUL_BIAS=0 \
READOUT_TARGETED=1 READOUT_INSTR_STYLE=pyrepr SHIM_STAGGER=1 \
python openjev-flash-9b/helper/shim.py --host 127.0.0.1 --port 3000

Then POST /v1/systemone as on the main card. curl -s http://127.0.0.1:3000/v1/version should show T: 1.07, noul_t: 1.074766, noul_bias: 0, targeted readout and instr_style: "pyrepr"; always pass the five READOUT_* settings above. FP8_CONVERSION_RECEIPT.json records the source checkpoint and the conversion method.

Licence

Weights: CC BY-NC 4.0 for research and non-commercial use, with attribution, the same as the main repository. For a commercial licence, email support@loopai.com. The licence texts and the base model's attribution are in LICENSE, LICENSE-APACHE-2.0 and NOTICE.

OpenJev is an independent project, not affiliated with TypeSafe; Jev is their product.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for openjev/OpenJev-Flash-9B-FP8

Finetuned
Qwen/Qwen3.5-9B
Quantized
(4)
this model