Instructions to use Remek/basal-1.5-mini-MLX-fp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Remek/basal-1.5-mini-MLX-fp4 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Remek/basal-1.5-mini-MLX-fp4") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use Remek/basal-1.5-mini-MLX-fp4 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Remek/basal-1.5-mini-MLX-fp4"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Remek/basal-1.5-mini-MLX-fp4" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Remek/basal-1.5-mini-MLX-fp4", "messages": [ {"role": "user", "content": "Hello"} ] }' - Atomic Chat
basal-1.5-mini-MLX-fp4
4-bit (nvfp4) MLX version of basal-1.5-mini (1.5B) for Apple Silicon. Same
model, same prompt format and the same typed answers (choice, noul, score) with a probability for every option;
see the main card for the description, benchmarks, limitations and license.
âš Quality warning. The 4-bit MLX port of basal-1.5-mini loses about 3.6 accuracy points against bf16 (0.882 vs 0.918) and its confidence thresholds are not certified. For mini, use
-MLX-8bitor GGUF (Q8_0/Q4_K_M), whose accuracy stays within noise of bf16.
Agreement with the bf16 model: 0.908; accuracy change −0.036 (0.882 vs 0.918) (500 held-out Polish development decisions, both option orders, calibrated, Apple M4 Pro; agreement = same top option as the reference (the same weights in PyTorch bf16; for basal-1.5-max, GGUF Q8_0 with llama.cpp); latency = median per decision (one question, both option orders) through the basal server path). Latency: 197 ms per decision on Apple M4 Pro, 24 GB. Memory: 1.0 GB. Accuracy here is on a 500-decision sample, to check that the port matches its reference; to compare model sizes, see the benchmark table of the main card.
Weights: MLX nvfp4 decoder layers (4-bit floating point E2M1, group 16, FP8 scales); the token embeddings and the output head stay in bf16, because the decision is read from the
output head's logits of the option letters.
Files: the MLX weights and config, tokenizer and chat template, CALIBRATION.json and CALIBRATION.mlx.json (this port's temperatures, used automatically by --mode mlx), basal.json.
Quick start
uv venv --python 3.12 ~/basal-mlx && source ~/basal-mlx/bin/activate
uv pip install "basal[mlx] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-mini-MLX-fp4 --mode mlx --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
"questions": {"dept": {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
"criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}}}}'
The basal server renders the model's own prompt, asks both option orders, applies the calibration and supports
multi, act and "facts": "auto"; the state of a request is computed once into a KV cache shared by all its
questions. Evidence spans are not available with MLX (use the bf16 model with the CUDA engine).
Engines
The status column refers to basal-1.5 (4.5B); basal-1.5-mini uses the same engines and API, and its own measurements are listed where available.
| engine | basal-serve mode |
hardware | what it is for | quick start | status (basal-1.5, 4.5B) |
|---|---|---|---|---|---|
| basal engine (PyTorch) | fast (also fast-nocompile, fp8, nvfp4); eager --device mps|cpu |
NVIDIA GPUs (sm80+); Apple Silicon (MPS) and CPU with eager |
the primary server: packed requests (SOAM), evidence spans, fp8 on Hopper and Blackwell |
basal-serve --model Remek/basal-1.5-mini --mode fast |
verified: served check on one H100, 13.7 ms p50 per question over HTTP, 64.4 decisions/s with 32 clients |
| vLLM | vllm |
NVIDIA GPUs | an alternative CUDA server | basal-serve --model Remek/basal-1.5-mini --mode vllm (extra basal[vllm]) |
verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.994 with the fp32 engine, accuracy 0.928 vs 0.930, 147 decisions/s |
| SGLang | sglang |
NVIDIA GPUs | an alternative CUDA server (RadixAttention prefix sharing) | basal-serve --model Remek/basal-1.5-mini --mode sglang (own environment: sglang[srt]==0.5.21, then basal with --no-deps) |
verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.995 with the fp32 engine, accuracy 0.925 vs 0.930, 104.6 decisions/s; on Werdykt the same answer as the basal engine on 98.3% of items (97.0% for states over 8k tokens) |
| MLX | mlx |
Apple Silicon | Macs: -MLX-8bit (recommended) or -MLX-fp4 (half the memory) |
basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx (extra basal[mlx]) |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.992 (-MLX-8bit) / 0.970 (-MLX-fp4), 587 / 581 ms per question |
| Ollama | ollama |
CPU, Apple Silicon, consumer GPUs | local and desktop use with Ollama (-GGUF) |
1. hf download Remek/basal-1.5-mini-GGUF --local-dir basal-1.5-mini-GGUF 2. in that folder: ollama create basal-1.5-mini:q8_0 -f Modelfile.Q8_0 3. basal-serve --model ./basal-1.5-mini-GGUF --mode ollama --ollama-model basal-1.5-mini:q8_0 |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 518 / 542 ms per question |
| llama.cpp | llamacpp |
CPU, Apple Silicon, consumer GPUs | llama-server with the -GGUF files; the engine's token ids are sent as is |
1. llama-server -m basal-1.5-mini-GGUF/basal-1.5-mini-Q8_0.gguf -c 4096 --port 8080 2. basal-serve --model ./basal-1.5-mini-GGUF --mode llamacpp |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 427 / 446 ms per question |
basal-serve is the decision layer in every row. vLLM, SGLang, MLX, Ollama and llama.cpp run the weights; the typed answers (a calibrated probability for each of your options, both option orders, the state shared by several questions, multi, act and facts) come from basal-serve --mode <engine> in front of them, with the same HTTP API in every mode. A plain ollama run or llama-cli gives free text only. Evidence spans need the PyTorch engine (fast, eager); vLLM and SGLang start with a 4,096-token context (longer states are refused); raise it with --max-len (e.g. 32768 for long documents). Install and first request for each engine: quick start.
Calibration
CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. This port's probabilities are flatter than the bf16 model's, so it ships CALIBRATION.mlx.json (used automatically by --mode mlx): the bf16 temperatures × 0.94 for choice and × 0.84 for yes/no questions, fitted on separate development decisions (expected calibration error 0.059 → 0.046). Its 1% and 5% confidence thresholds are copied from the bf16 model and are not certified for this port: the temperatures were refitted, so the bf16 model's coverage and error figures do not carry over. Refit the thresholds on your own labelled requests before automating decisions with them.
License
Apache-2.0, like basal-1.5-mini.
- Downloads last month
- 27
4-bit
Model tree for Remek/basal-1.5-mini-MLX-fp4
Base model
speakleash/Bielik-1.5B-v3