OpenJev Flash 9B, FP8 checkpoint
The FP8 checkpoint of OpenJev Flash 9B: the same weights stored as FP8 (e4m3, 128×128 weight blocks, dynamic activations), 11.9 GB instead of 18.8 GB.
On par with Cloudflare's Clef-Flash and ahead of Kev-9B and Nimble 9B on JevBench. The results, the API and the use cases are on the main card.
- Drop-in for vLLM: the main card's command and helper, pointed at this repository. The weights are already FP8, so nothing is quantized at load time.
- Matches the 16-bit model: 80.0% on a fixed 1,789-question check, the same as the 16-bit weights, and 79.2% on the frozen 10,000 text questions against 79.4% for the 16-bit weights served in FP8.
- 37% smaller download: 11.9 GB instead of 18.8 GB.
Measured
One GPU, vLLM 0.29. The 10,000 questions go through the OpenJev helper with the card's calibration; the 1,789-question check reads the letter probabilities directly.
| test | 16-bit weights | 16-bit, served in FP8 | this FP8 file |
|---|---|---|---|
| 10,000 text questions, 34 public sources | — | 79.4% | 79.2% |
| fixed 1,789-question check | 80.0% | 79.7% | 80.0% |
Run it
pip install "vllm==0.29.0" "openai==3.26.0" "httpx==0.28.1"
hf download openjev/OpenJev-Flash-9B helper/shim.py --local-dir openjev-flash-9b
Terminal 1, the model (downloads the weights on first start):
vllm serve openjev/OpenJev-Flash-9B-FP8 --host 127.0.0.1 --served-model-name qwen --port 8000 \
--enable-prefix-caching --max-model-len 16384 --gpu-memory-utilization 0.90 \
--limit-mm-per-prompt '{"image":1}' --trust-remote-code --max-num-seqs 16 \
--max-logprobs 64 --gdn-prefill-backend triton --quantization fp8
Terminal 2, the decision API in front of it:
VLLM=http://127.0.0.1:8000/v1 TOKENIZER=openjev/OpenJev-Flash-9B-FP8 \
READOUT_T=1.07 READOUT_NOUL_T=1.074766 READOUT_NOUL_BIAS=0 \
READOUT_TARGETED=1 READOUT_INSTR_STYLE=pyrepr SHIM_STAGGER=1 \
python openjev-flash-9b/helper/shim.py --host 127.0.0.1 --port 3000
Then POST /v1/systemone as on the main card. curl -s http://127.0.0.1:3000/v1/version should show T: 1.07, noul_t: 1.074766, noul_bias: 0, targeted readout and instr_style: "pyrepr"; always pass the five READOUT_* settings above. FP8_CONVERSION_RECEIPT.json records the source checkpoint and the conversion method.
Licence
Weights: CC BY-NC 4.0 for research and non-commercial use, with attribution, the same as the main repository. For a commercial licence, email support@loopai.com. The licence texts and the base model's attribution are in LICENSE, LICENSE-APACHE-2.0 and NOTICE.
OpenJev is an independent project, not affiliated with TypeSafe; Jev is their product.
- Downloads last month
- -
