Atlas3D/JEV-27B-VL-FP8
An FP8 (W8A8, dynamic per-token activations) checkpoint of
autotrust/JEV-27B-VL, made with llm-compressor
(FP8_DYNAMIC, compressed-tensors format). Weights are about 34.3 GiB, against about 53 GiB
for the bf16 original. It needs a GPU with FP8 support (Ada, Hopper or Blackwell).
What is quantized: the language model's Linear layers in its full-attention blocks and MLPs.
What stays bf16: the vision tower, the linear-attention blocks, the embeddings, lm_head, MTP, and
the JEV System 1 decision adapter (adapter_vllm/, unchanged).
Evaluation (System 1 /v1/decide)
Same harness and states as the bf16 model, compared one decision at a time:
| result | |
|---|---|
| Top choice agrees with bf16 | 299/308 (97.1%) |
| Largest probability change on any option | 0.061 |
| Sanity scenarios | sanity 6/6 |
This is a narrow test of the decision head. System 2 (free-text generation) quality was not separately benchmarked.
Serving
Serve it the same way as the original. serve_decide.py is included, and vLLM reads the quantization
from config.json:
python serve_decide.py --model Atlas3D/JEV-27B-VL-FP8 --enable-lora --max-lora-rank 32 \
--lora-modules jev-decision=<path>/adapter_vllm --logprobs-mode processed_logprobs \
--enable-prefix-caching --mamba-cache-mode align --max-num-seqs 8 --trust-request-chat-template
See the upstream card for the System 1 API, intended use and limitations, all of which apply unchanged.
License
Apache-2.0, as with the original. This is a modified (quantized) version of autotrust/JEV-27B-VL,
which is itself based on Qwen/Qwen3.8-27B. LICENSE is included unchanged.
- Downloads last month
- 17