Instructions to use suhaas-teja/Dev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use suhaas-teja/Dev-4B with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download suhaas-teja/Dev-4B --local-dir Dev-4B
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Dev-4B
A small model that answers questions about a document instantly, with a confidence score, and stops to think step by step only when a question is too hard for a quick answer. It runs on your own computer, including a 16 GB Mac.
Dev-4B is Qwen3-4B-Instruct-2507 plus four small trained add-ons (133 MB). The base model is never changed: with the add-ons switched off, it generates exactly what Qwen3-4B generates.
What it does
1. Quick decisions with honest confidence. Give it a document and questions with fixed answers: yes/no, pick one of up to 16 options, or a score on up to 6 levels. This is Kev-4B's example ticket, answered by Dev-4B on a Mac:
"I was charged twice for order 1182. Please refund one of the charges."
| Question | Dev-4B's answer |
|---|---|
| Which team should handle this? | billing 0.99, returns 0.005, shipping 0.005 |
| Does this need a reply today? | yes 0.42, no 0.58: the ticket gives no deadline, and the model says it isn't sure |
Both questions share one reading of the document and come back together in about a second. On the kinds of task it was trained on, the confidence is calibrated: when it says 90%, it is right about 90% of the time.
2. Thinking when a quick answer isn't enough. Decision-only models stop at the quick answer, even on questions that need arithmetic. Dev-4B has a small built-in checker, the router, that notices when the quick answer is likely wrong; the same model then reasons step by step from the document it has already read:
"I was charged twice for order 1182, $48.00 each time. Please refund one charge, and as compensation also refund 25% of the other. Store policy: any refund above $50.00 needs a manager to approve it." Does this refund need a manager to approve it?
- Quick answer: yes 0.64. Unsure.
- Router risk 0.83, above its 0.516 threshold, so it escalates.
- Reasoning, from the same reading of the document: "25% of $48.00 = $12.00 … total refund = $48.00 + $12.00 = $60.00 … Since $60.00 > $50.00, a manager approval is required. Answer: Yes"
How well it works
All numbers are from 7,400 frozen test questions, built and hashed before any training.
| On 7,100 scored test questions | Accuracy | Time per question (H100 GPU, modelled) |
|---|---|---|
| Always answer in one quick pass | 76.3% | 0.16 s |
| Always reason step by step | 78.0% | 10.2 s |
| Dev-4B: quick, reasoning only when flagged | 81.8% | 1.37 s |
| One quick pass | Zero-shot base | Kev-4B | Dev-4B |
|---|---|---|---|
| Accuracy, task types it was trained on | 76.6% | 75.8% | 83.4% |
| Accuracy, task types never seen in training | 74.8% | 76.6% | 74.2% |
| Calibration error, trained task types (lower is better) | 0.213 | 0.120 | 0.040 |
| Calibration error, unseen task types | 0.227 | 0.115 | 0.139 |
| Answers that change when options are shuffled | 7.4% | 5.1% | 3.8% |
- The gain from reasoning is concentrated on arithmetic and counting: 12% → 99% and 47% → 93% on those test sets. On 250 arithmetic questions, Kev-4B answers 50.4% in one pass.
- On a 16 GB M1 Pro (MLX, 8-bit), five questions about one ~600-token document take about 3 s, and the one-pass model peaks at about 6 GB of memory.
- Across three training seeds, accuracy on trained task types is 83.4% ± 0.1.
Limits
- Unfamiliar tasks: on task types it never trained on, it is no more accurate than its base model, and its confidence is less reliable than Kev-4B's. Recalibrate before relying on its probabilities for a new task.
- Reasoning is a 4B model's reasoning. Escalation fixes arithmetic and counting well, multi-hop lookups much less (56% → 64%). It is far weaker than large hosted models.
- English only, documents up to about 4,000 characters in training.
- Not for consequential decisions about people (legal, medical, financial, employment) without human review.
Run it
You need Python 3.12, the Hugging Face command line tool (pip install -U huggingface_hub) and either an Apple-silicon Mac with 16 GB or an NVIDIA GPU with about 10 GB. The add-ons are in this repository; the base model downloads from Qwen's repository at a pinned revision.
On a Mac (uses the ready-made 8-bit build, suhaas-teja/Dev-4B-MLX-8bit, about 4 GB):
hf download suhaas-teja/Dev-4B --local-dir Dev-4B && cd Dev-4B
hf download suhaas-teja/Dev-4B-MLX-8bit --local-dir mlx-8bit
pip install -e ".[serve,mac]"
python -m serve.server_mlx --model mlx-8bit --artifacts .
On an NVIDIA GPU (downloads the 8 GB base on first start):
hf download suhaas-teja/Dev-4B --local-dir Dev-4B && cd Dev-4B
pip install -e ".[serve]"
python -m serve.server --artifacts .
Then open http://127.0.0.1:8008 for a playground in your browser, or call the API:
curl -s localhost:8008/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice for order 1182. Please refund one of the charges.",
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}}}}'
The server has three modes: quick answers (/v1/systemone), plain chat with the base model (/v1/chat/completions) and automatic (/v1/auto). Requests and responses are in docs/api.md.
It does not run in Ollama or LM Studio. Ollama no longer accepts LoRA adapters at all, and LM Studio runs only single merged models; neither can run the decision head or the router. A llama.cpp version works in testing (179 of 180 test decisions match the Mac build) but is not included here yet.
How it works
- An adapter (LoRA, rank 16, on every attention and MLP projection) that is switched off while the document is read and on for the question. The model's memory of the document is therefore exactly the base model's, so the quick answer and the reasoning can share it.
- A decision head that turns the last token's hidden state into one score per allowed answer (24 slots: yes/no, levels 0–5, letters A–P).
- Temperatures, one per question kind (yes/no 0.59, choice 0.82, score 0.89), that make those scores honest probabilities.
- A router, a two-layer network of width 8, that reads the same hidden state plus three confidence features and predicts when the quick answer is likely wrong.
| File | What it is |
|---|---|
adapter_and_head.safetensors |
The adapter (504 tensors) and the decision head |
router.safetensors |
The router |
temperatures.json, operating_point.json |
The calibration temperatures and the router's threshold (0.516) |
dmm/, serve/, train/train_decision.py, configs/ |
The code needed to run it (the project's working name is dmm) |
Training: 30,000 records (35,447 rows with augmented copies) from BoolQ, PAWS, STS-B, HelpSteer2, ARC, CommonsenseQA, QASC, HellaSwag and a synthetic table-lookup task, with targets from Qwen3-30B-A3B-Instruct-2507 where it agreed with the gold label. Every training record was checked against every test item for duplicates. 1,330 steps, about 34 minutes on one H100.
Credits
Dev-4B is an independent project, not affiliated with TypeSafe AI. Its question format follows TypeSafe's System One API, as Kev-4B's does. The decision head follows the design of JEV-27B. Switching the adapter off on the document follows Activated LoRA and Efficient Reasoning on the Edge.
Licence
- Code (
dmm/,serve/,train/,configs/): Apache-2.0, seeLICENSE. - Weights (
*.safetensors,*.json): CC BY-NC-SA 4.0, non-commercial, seeLICENSE-WEIGHTS.md. The training data includes sources with share-alike terms (BoolQ CC-BY-SA-3.0, ARC CC-BY-SA-4.0) and research-only terms (parts of STS-B), so the weights are released for non-commercial use. - Base model: Qwen3-4B-Instruct-2507 is Apache-2.0 and is not included here.
| Training source | Licence |
|---|---|
| BoolQ | CC-BY-SA-3.0 |
| PAWS | free for any purpose with attribution to Google |
| STS-B | mixed per sub-corpus, research use |
| HelpSteer2 | CC-BY-4.0 |
| ARC | CC-BY-SA-4.0 |
| CommonsenseQA | MIT |
| QASC | CC-BY-4.0 |
| HellaSwag | MIT |
| Table lookup | generated by this project |
Quantized
Model tree for suhaas-teja/Dev-4B
Base model
Qwen/Qwen3-4B-Instruct-2507