basal-1.5-mini-MLX-fp4

GitHub Website Collection

4-bit (nvfp4) MLX version of basal-1.5-mini (1.5B) for Apple Silicon. Same model, same prompt format and the same typed answers (choice, noul, score) with a probability for every option; see the main card for the description, benchmarks, limitations and license.

âš  Quality warning. The 4-bit MLX port of basal-1.5-mini loses about 3.6 accuracy points against bf16 (0.882 vs 0.918) and its confidence thresholds are not certified. For mini, use -MLX-8bit or GGUF (Q8_0 / Q4_K_M), whose accuracy stays within noise of bf16.

Agreement with the bf16 model: 0.908; accuracy change −0.036 (0.882 vs 0.918) (500 held-out Polish development decisions, both option orders, calibrated, Apple M4 Pro; agreement = same top option as the reference (the same weights in PyTorch bf16; for basal-1.5-max, GGUF Q8_0 with llama.cpp); latency = median per decision (one question, both option orders) through the basal server path). Latency: 197 ms per decision on Apple M4 Pro, 24 GB. Memory: 1.0 GB. Accuracy here is on a 500-decision sample, to check that the port matches its reference; to compare model sizes, see the benchmark table of the main card.

Weights: MLX nvfp4 decoder layers (4-bit floating point E2M1, group 16, FP8 scales); the token embeddings and the output head stay in bf16, because the decision is read from the output head's logits of the option letters.

Files: the MLX weights and config, tokenizer and chat template, CALIBRATION.json and CALIBRATION.mlx.json (this port's temperatures, used automatically by --mode mlx), basal.json.

Quick start

uv venv --python 3.12 ~/basal-mlx && source ~/basal-mlx/bin/activate
uv pip install "basal[mlx] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-mini-MLX-fp4 --mode mlx --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
  "questions": {"dept": {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
    "criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}}}}'

The basal server renders the model's own prompt, asks both option orders, applies the calibration and supports multi, act and "facts": "auto"; the state of a request is computed once into a KV cache shared by all its questions. Evidence spans are not available with MLX (use the bf16 model with the CUDA engine).

Engines

The status column refers to basal-1.5 (4.5B); basal-1.5-mini uses the same engines and API, and its own measurements are listed where available.

engine basal-serve mode hardware what it is for quick start status (basal-1.5, 4.5B)
basal engine (PyTorch) fast (also fast-nocompile, fp8, nvfp4); eager --device mps|cpu NVIDIA GPUs (sm80+); Apple Silicon (MPS) and CPU with eager the primary server: packed requests (SOAM), evidence spans, fp8 on Hopper and Blackwell basal-serve --model Remek/basal-1.5-mini --mode fast verified: served check on one H100, 13.7 ms p50 per question over HTTP, 64.4 decisions/s with 32 clients
vLLM vllm NVIDIA GPUs an alternative CUDA server basal-serve --model Remek/basal-1.5-mini --mode vllm (extra basal[vllm]) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.994 with the fp32 engine, accuracy 0.928 vs 0.930, 147 decisions/s
SGLang sglang NVIDIA GPUs an alternative CUDA server (RadixAttention prefix sharing) basal-serve --model Remek/basal-1.5-mini --mode sglang (own environment: sglang[srt]==0.5.21, then basal with --no-deps) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.995 with the fp32 engine, accuracy 0.925 vs 0.930, 104.6 decisions/s; on Werdykt the same answer as the basal engine on 98.3% of items (97.0% for states over 8k tokens)
MLX mlx Apple Silicon Macs: -MLX-8bit (recommended) or -MLX-fp4 (half the memory) basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx (extra basal[mlx]) verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.992 (-MLX-8bit) / 0.970 (-MLX-fp4), 587 / 581 ms per question
Ollama ollama CPU, Apple Silicon, consumer GPUs local and desktop use with Ollama (-GGUF) 1. hf download Remek/basal-1.5-mini-GGUF --local-dir basal-1.5-mini-GGUF 2. in that folder: ollama create basal-1.5-mini:q8_0 -f Modelfile.Q8_0 3. basal-serve --model ./basal-1.5-mini-GGUF --mode ollama --ollama-model basal-1.5-mini:q8_0 verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 518 / 542 ms per question
llama.cpp llamacpp CPU, Apple Silicon, consumer GPUs llama-server with the -GGUF files; the engine's token ids are sent as is 1. llama-server -m basal-1.5-mini-GGUF/basal-1.5-mini-Q8_0.gguf -c 4096 --port 8080 2. basal-serve --model ./basal-1.5-mini-GGUF --mode llamacpp verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 427 / 446 ms per question

basal-serve is the decision layer in every row. vLLM, SGLang, MLX, Ollama and llama.cpp run the weights; the typed answers (a calibrated probability for each of your options, both option orders, the state shared by several questions, multi, act and facts) come from basal-serve --mode <engine> in front of them, with the same HTTP API in every mode. A plain ollama run or llama-cli gives free text only. Evidence spans need the PyTorch engine (fast, eager); vLLM and SGLang start with a 4,096-token context (longer states are refused); raise it with --max-len (e.g. 32768 for long documents). Install and first request for each engine: quick start.

Calibration

CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. This port's probabilities are flatter than the bf16 model's, so it ships CALIBRATION.mlx.json (used automatically by --mode mlx): the bf16 temperatures × 0.94 for choice and × 0.84 for yes/no questions, fitted on separate development decisions (expected calibration error 0.059 → 0.046). Its 1% and 5% confidence thresholds are copied from the bf16 model and are not certified for this port: the temperatures were refitted, so the bf16 model's coverage and error figures do not carry over. Refit the thresholds on your own labelled requests before automating decisions with them.

License

Apache-2.0, like basal-1.5-mini.

Downloads last month
27
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Remek/basal-1.5-mini-MLX-fp4

Quantized
(5)
this model

Collection including Remek/basal-1.5-mini-MLX-fp4