basal-1.5-mini-MLX-8bit

GitHub Website Collection

8-bit MLX version of basal-1.5-mini (1.5B) for Apple Silicon. Same model, same prompt format and the same typed answers (choice, noul, score) with a probability for every option; see the main card for the description, benchmarks, limitations and license.

Agreement with the bf16 model: 0.998; accuracy change −0.002 (0.916 vs 0.918) (500 held-out Polish development decisions, both option orders, calibrated, Apple M4 Pro; agreement = same top option as the reference (the same weights in PyTorch bf16; for basal-1.5-max, GGUF Q8_0 with llama.cpp); latency = median per decision (one question, both option orders) through the basal server path). Latency: 196 ms per decision on Apple M4 Pro, 24 GB. Memory: 1.8 GB. Accuracy here is on a 500-decision sample, to check that the port matches its reference; to compare model sizes, see the benchmark table of the main card.

Weights: MLX affine 8-bit decoder layers (group 64, a bf16 scale and bias per group); the token embeddings and the output head stay in bf16, because the decision is read from the output head's logits of the option letters. The recommended Apple Silicon port; -MLX-fp4 needs about half the memory.

Files: the MLX weights and config, tokenizer and chat template, CALIBRATION.json, basal.json.

Quick start

uv venv --python 3.12 ~/basal-mlx && source ~/basal-mlx/bin/activate
uv pip install "basal[mlx] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
  "questions": {"dept": {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
    "criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}}}}'

The basal server renders the model's own prompt, asks both option orders, applies the calibration and supports multi, act and "facts": "auto"; the state of a request is computed once into a KV cache shared by all its questions. Evidence spans are not available with MLX (use the bf16 model with the CUDA engine).

Engines

The status column refers to basal-1.5 (4.5B); basal-1.5-mini uses the same engines and API, and its own measurements are listed where available.

engine basal-serve mode hardware what it is for quick start status (basal-1.5, 4.5B)
basal engine (PyTorch) fast (also fast-nocompile, fp8, nvfp4); eager --device mps|cpu NVIDIA GPUs (sm80+); Apple Silicon (MPS) and CPU with eager the primary server: packed requests (SOAM), evidence spans, fp8 on Hopper and Blackwell basal-serve --model Remek/basal-1.5-mini --mode fast verified: served check on one H100, 13.7 ms p50 per question over HTTP, 64.4 decisions/s with 32 clients
vLLM vllm NVIDIA GPUs an alternative CUDA server basal-serve --model Remek/basal-1.5-mini --mode vllm (extra basal[vllm]) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.994 with the fp32 engine, accuracy 0.928 vs 0.930, 147 decisions/s
SGLang sglang NVIDIA GPUs an alternative CUDA server (RadixAttention prefix sharing) basal-serve --model Remek/basal-1.5-mini --mode sglang (own environment: sglang[srt]==0.5.21, then basal with --no-deps) verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.995 with the fp32 engine, accuracy 0.925 vs 0.930, 104.6 decisions/s; on Werdykt the same answer as the basal engine on 98.3% of items (97.0% for states over 8k tokens)
MLX mlx Apple Silicon Macs: -MLX-8bit (recommended) or -MLX-fp4 (half the memory) basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx (extra basal[mlx]) verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.992 (-MLX-8bit) / 0.970 (-MLX-fp4), 587 / 581 ms per question
Ollama ollama CPU, Apple Silicon, consumer GPUs local and desktop use with Ollama (-GGUF) 1. hf download Remek/basal-1.5-mini-GGUF --local-dir basal-1.5-mini-GGUF 2. in that folder: ollama create basal-1.5-mini:q8_0 -f Modelfile.Q8_0 3. basal-serve --model ./basal-1.5-mini-GGUF --mode ollama --ollama-model basal-1.5-mini:q8_0 verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 518 / 542 ms per question
llama.cpp llamacpp CPU, Apple Silicon, consumer GPUs llama-server with the -GGUF files; the engine's token ids are sent as is 1. llama-server -m basal-1.5-mini-GGUF/basal-1.5-mini-Q8_0.gguf -c 4096 --port 8080 2. basal-serve --model ./basal-1.5-mini-GGUF --mode llamacpp verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 427 / 446 ms per question

basal-serve is the decision layer in every row. vLLM, SGLang, MLX, Ollama and llama.cpp run the weights; the typed answers (a calibrated probability for each of your options, both option orders, the state shared by several questions, multi, act and facts) come from basal-serve --mode <engine> in front of them, with the same HTTP API in every mode. A plain ollama run or llama-cli gives free text only. Evidence spans need the PyTorch engine (fast, eager); vLLM and SGLang start with a 4,096-token context (longer states are refused); raise it with --max-len (e.g. 32768 for long documents). Install and first request for each engine: quick start.

Calibration

CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. Checked on this port: the temperature that fits it best is 1.00 times the bf16 one, so no separate calibration file is needed. The confidence thresholds are not certified for this port: refit them on your own labelled requests before automating decisions with them.

License

Apache-2.0, like basal-1.5-mini.

Downloads last month
33
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Remek/basal-1.5-mini-MLX-8bit

Quantized
(5)
this model

Collection including Remek/basal-1.5-mini-MLX-8bit