Text Classification
MLX
English
decision-model
calibration
lora
activated-lora
adaptive-reasoning
on-device
Instructions to use suhaas-teja/Dev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use suhaas-teja/Dev-4B with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download suhaas-teja/Dev-4B --local-dir Dev-4B
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
File size: 6,128 Bytes
a69969c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 | # API
One server process answers three kinds of request. The GPU server (`serve/server.py`, PyTorch) and the Mac server (`serve/server_mlx.py`, MLX) share the same routes and formats.
| Route | What it does |
| --- | --- |
| `POST /v1/systemone` | Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag |
| `POST /v1/chat/completions` | Plain chat with the base model (the adapter is off) |
| `POST /v1/auto` | One question: decide, and if the router flags it, answer by generating from the same document cache |
Start a server:
```bash
# From the downloaded Dev-4B folder (see README.md)
python -m serve.server --artifacts . # NVIDIA GPU
python -m serve.server_mlx --model mlx-8bit --artifacts . # Mac (Dev-4B-MLX-8bit in mlx-8bit)
```
The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly.
## Decisions: `POST /v1/systemone`
The *state* is the document. Each question is one of three types:
| Type | Options | Answer fields |
| --- | --- | --- |
| `noul` | yes / no | `noul`: probability of yes |
| `choice` | `criteria`: 1 to 16 named options, each with an optional description | `choice`: the most likely option name |
| `score` | `criteria`: 2 to 6 ordered level descriptions, lowest first | `score`: expected level, 0 = first |
Every answer also has `probabilities`, a `router_score` (0 to 1, higher means a single pass is more likely to be wrong) and `needs_generation` (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state.
Request:
```json
{
"state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"tone": {"type": "score", "instructions": "How upset is the customer?",
"criteria": ["calm", "annoyed", "angry"]}
}
}
```
Response (trimmed):
```json
{
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009},
"router_score": 0.246, "needs_generation": false},
"urgent": {"type": "noul", "noul": 0.615,
"probabilities": {"false": 0.385, "true": 0.615},
"router_score": 0.799, "needs_generation": true},
"tone": {"type": "score", "score": 1.46,
"probabilities": {"0": 0.038, "1": 0.462, "2": 0.500},
"router_score": 0.363, "needs_generation": false}
},
"latency_ms": 861.5
}
```
Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry".
Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns `422` with a message.
## Generation: `POST /v1/chat/completions`
The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes `messages` (OpenAI-style roles) and `max_tokens`. Decoding is greedy.
```json
{"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40}
```
```json
{
"object": "chat.completion",
"choices": [{"index": 0, "finish_reason": "stop",
"message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}],
"tokens": [32, 20965, 374, "..."],
"usage": {"completion_tokens": 30}
}
```
`tokens` carries the raw token ids, which is how the identity checks compare outputs.
## Decide, then generate if needed: `POST /v1/auto`
One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (`state_prefills` is always 1).
```json
{
"state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.",
"name": "count",
"question": {"type": "choice", "instructions": "How many stamps does Tara have now?",
"criteria": {"12": null, "13": null, "24": null, "14": null}},
"max_tokens": 300
}
```
```json
{
"decision": {"type": "choice", "choice": "14",
"probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587},
"router_score": 0.855, "needs_generation": true},
"escalated": true,
"state_prefills": 1,
"generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A",
"generation_tokens": ["126 ids"]
}
```
Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When `escalated` is false the response has only `decision`.
## Speed
| | First token after escalation, cache reuse | Encoding the document again |
| --- | --- | --- |
| Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s |
| Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s |
| H100, bf16, 4,091-token document | 33 ms | 98 ms |
| H100, bf16, 257-token document | 32 ms | 29 ms |
A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100.
## Limits
- One request on the model at a time; requests queue.
- English, and documents up to about 4,000 characters were trained on.
- Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what `auto` is for.
|