basal-1.5-max-FP8

GitHub Website Collection

FP8 (NVIDIA Model Optimizer) checkpoint of basal-1.5-max (11B) for vLLM on NVIDIA GPUs. Same model, same prompt format and the same typed answers (choice, noul, score) with a probability for every option; see the main card for the description, benchmarks, limitations and license.

Agreement with the bf16 model (fp32 reference): 0.989; accuracy 0.941 (−0.003; bf16 0.944) (1,000 development decisions, both option orders, vLLM 0.30.0 on one RTX PRO 6000 Blackwell; reference: the bf16 weights in fp32). Speed: 24.5 ms per decision at batch 1, 76.4 decisions/s batched, on one RTX PRO 6000 Blackwell (vLLM 0.30.0). Size: 11.4 GB.

What it is

Post-training quantisation with NVIDIA Model Optimizer (FP8_DEFAULT_CFG), calibrated on basal-1.5 calibration decisions: FP8 (E4M3) weights and activations in the decoder layers. The token embeddings and the output head stay in bf16, because the decision is read from the output head's logits. It is a standard ModelOpt Hugging Face export.

Hardware: NVIDIA Hopper (H100, H200), Ada (RTX 40xx, L40S) and Blackwell (B200, B300, RTX 50xx, RTX PRO 6000, DGX Spark).

Run

vLLM brings its own torch, so use a separate environment:

uv pip install "basal[vllm] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-max-FP8 --mode vllm --port 8000

--mode vllm starts with a 4,096-token context and refuses longer states; raise it with --max-len (e.g. 32768 for long documents). The request and response are the same as for the bf16 model. The basal engine's own modes do not load this checkpoint; for on-the-fly FP8 in the basal engine, use the bf16 repository with --mode fp8. Evidence spans need the basal engine.

Calibration

CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. The thresholds are not validated for this checkpoint: refit them on your own labelled requests before automating decisions with them.

License

Apache-2.0, like basal-1.5-max.

Downloads last month
13
Safetensors
Model size
11B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Remek/basal-1.5-max-FP8

Collection including Remek/basal-1.5-max-FP8