Instructions to use Remek/basal-1.5-mini-MLX-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Remek/basal-1.5-mini-MLX-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Remek/basal-1.5-mini-MLX-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use Remek/basal-1.5-mini-MLX-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Remek/basal-1.5-mini-MLX-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Remek/basal-1.5-mini-MLX-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Remek/basal-1.5-mini-MLX-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Atomic Chat
basal-1.5-mini-MLX-8bit
8-bit MLX version of basal-1.5-mini (1.5B) for Apple Silicon. Same
model, same prompt format and the same typed answers (choice, noul, score) with a probability for every option;
see the main card for the description, benchmarks, limitations and license.
Agreement with the bf16 model: 0.998; accuracy change −0.002 (0.916 vs 0.918) (500 held-out Polish development decisions, both option orders, calibrated, Apple M4 Pro; agreement = same top option as the reference (the same weights in PyTorch bf16; for basal-1.5-max, GGUF Q8_0 with llama.cpp); latency = median per decision (one question, both option orders) through the basal server path). Latency: 196 ms per decision on Apple M4 Pro, 24 GB. Memory: 1.8 GB. Accuracy here is on a 500-decision sample, to check that the port matches its reference; to compare model sizes, see the benchmark table of the main card.
Weights: MLX affine 8-bit decoder layers (group 64, a bf16 scale and bias per group); the token embeddings and the output head stay in bf16, because the decision is read from the
output head's logits of the option letters. The recommended Apple Silicon port; -MLX-fp4 needs about half the memory.
Files: the MLX weights and config, tokenizer and chat template, CALIBRATION.json, basal.json.
Quick start
uv venv --python 3.12 ~/basal-mlx && source ~/basal-mlx/bin/activate
uv pip install "basal[mlx] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "Klient: od wczoraj nie mogę zalogować się do bankowości internetowej, system pokazuje błąd hasła.",
"questions": {"dept": {"type": "choice", "instructions": "Do którego działu skierować zgłoszenie?",
"criteria": {"cards": "Reklamacje kart", "online": "Wsparcie bankowości elektronicznej", "loans": "Kredyty"}}}}'
The basal server renders the model's own prompt, asks both option orders, applies the calibration and supports
multi, act and "facts": "auto"; the state of a request is computed once into a KV cache shared by all its
questions. Evidence spans are not available with MLX (use the bf16 model with the CUDA engine).
Engines
The status column refers to basal-1.5 (4.5B); basal-1.5-mini uses the same engines and API, and its own measurements are listed where available.
| engine | basal-serve mode |
hardware | what it is for | quick start | status (basal-1.5, 4.5B) |
|---|---|---|---|---|---|
| basal engine (PyTorch) | fast (also fast-nocompile, fp8, nvfp4); eager --device mps|cpu |
NVIDIA GPUs (sm80+); Apple Silicon (MPS) and CPU with eager |
the primary server: packed requests (SOAM), evidence spans, fp8 on Hopper and Blackwell |
basal-serve --model Remek/basal-1.5-mini --mode fast |
verified: served check on one H100, 13.7 ms p50 per question over HTTP, 64.4 decisions/s with 32 clients |
| vLLM | vllm |
NVIDIA GPUs | an alternative CUDA server | basal-serve --model Remek/basal-1.5-mini --mode vllm (extra basal[vllm]) |
verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.994 with the fp32 engine, accuracy 0.928 vs 0.930, 147 decisions/s |
| SGLang | sglang |
NVIDIA GPUs | an alternative CUDA server (RadixAttention prefix sharing) | basal-serve --model Remek/basal-1.5-mini --mode sglang (own environment: sglang[srt]==0.5.21, then basal with --no-deps) |
verified offline (one H100, basal-bench, 1,000 development decisions): agreement 0.995 with the fp32 engine, accuracy 0.925 vs 0.930, 104.6 decisions/s; on Werdykt the same answer as the basal engine on 98.3% of items (97.0% for states over 8k tokens) |
| MLX | mlx |
Apple Silicon | Macs: -MLX-8bit (recommended) or -MLX-fp4 (half the memory) |
basal-serve --model Remek/basal-1.5-mini-MLX-8bit --mode mlx (extra basal[mlx]) |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.992 (-MLX-8bit) / 0.970 (-MLX-fp4), 587 / 581 ms per question |
| Ollama | ollama |
CPU, Apple Silicon, consumer GPUs | local and desktop use with Ollama (-GGUF) |
1. hf download Remek/basal-1.5-mini-GGUF --local-dir basal-1.5-mini-GGUF 2. in that folder: ollama create basal-1.5-mini:q8_0 -f Modelfile.Q8_0 3. basal-serve --model ./basal-1.5-mini-GGUF --mode ollama --ollama-model basal-1.5-mini:q8_0 |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 518 / 542 ms per question |
| llama.cpp | llamacpp |
CPU, Apple Silicon, consumer GPUs | llama-server with the -GGUF files; the engine's token ids are sent as is |
1. llama-server -m basal-1.5-mini-GGUF/basal-1.5-mini-Q8_0.gguf -c 4096 --port 8080 2. basal-serve --model ./basal-1.5-mini-GGUF --mode llamacpp |
verified on an Apple M4 Pro (500 development decisions): agreement with bf16 0.994 (Q8_0) / 0.980 (Q4_K_M), 427 / 446 ms per question |
basal-serve is the decision layer in every row. vLLM, SGLang, MLX, Ollama and llama.cpp run the weights; the typed answers (a calibrated probability for each of your options, both option orders, the state shared by several questions, multi, act and facts) come from basal-serve --mode <engine> in front of them, with the same HTTP API in every mode. A plain ollama run or llama-cli gives free text only. Evidence spans need the PyTorch engine (fast, eager); vLLM and SGLang start with a 4,096-token context (longer states are refused); raise it with --max-len (e.g. 32768 for long documents). Install and first request for each engine: quick start.
Calibration
CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. Checked on this port: the temperature that fits it best is 1.00 times the bf16 one, so no separate calibration file is needed. The confidence thresholds are not certified for this port: refit them on your own labelled requests before automating decisions with them.
License
Apache-2.0, like basal-1.5-mini.
- Downloads last month
- 33
8-bit
Model tree for Remek/basal-1.5-mini-MLX-8bit
Base model
speakleash/Bielik-1.5B-v3