Instructions to use suhaas-teja/Dev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use suhaas-teja/Dev-4B with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download suhaas-teja/Dev-4B --local-dir Dev-4B
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Download docs/api.md from suhaas-teja/Dev-4B: direct link, hf CLI and curl.
- Browser
- Download file 6.13 kB
-
https://huggingface.co/suhaas-teja/Dev-4B/resolve/main/docs/api.md
- Command line
-
hf download hf://suhaas-teja/Dev-4B/docs/api.md
-
curl -L -o api.md https://huggingface.co/suhaas-teja/Dev-4B/resolve/main/docs/api.md
API
One server process answers three kinds of request. The GPU server (serve/server.py, PyTorch) and the Mac server (serve/server_mlx.py, MLX) share the same routes and formats.
| Route | What it does |
|---|---|
POST /v1/systemone |
Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag |
POST /v1/chat/completions |
Plain chat with the base model (the adapter is off) |
POST /v1/auto |
One question: decide, and if the router flags it, answer by generating from the same document cache |
Start a server:
# From the downloaded Dev-4B folder (see README.md)
python -m serve.server --artifacts . # NVIDIA GPU
python -m serve.server_mlx --model mlx-8bit --artifacts . # Mac (Dev-4B-MLX-8bit in mlx-8bit)
The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly.
Decisions: POST /v1/systemone
The state is the document. Each question is one of three types:
| Type | Options | Answer fields |
|---|---|---|
noul |
yes / no | noul: probability of yes |
choice |
criteria: 1 to 16 named options, each with an optional description |
choice: the most likely option name |
score |
criteria: 2 to 6 ordered level descriptions, lowest first |
score: expected level, 0 = first |
Every answer also has probabilities, a router_score (0 to 1, higher means a single pass is more likely to be wrong) and needs_generation (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state.
Request:
{
"state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"tone": {"type": "score", "instructions": "How upset is the customer?",
"criteria": ["calm", "annoyed", "angry"]}
}
}
Response (trimmed):
{
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009},
"router_score": 0.246, "needs_generation": false},
"urgent": {"type": "noul", "noul": 0.615,
"probabilities": {"false": 0.385, "true": 0.615},
"router_score": 0.799, "needs_generation": true},
"tone": {"type": "score", "score": 1.46,
"probabilities": {"0": 0.038, "1": 0.462, "2": 0.500},
"router_score": 0.363, "needs_generation": false}
},
"latency_ms": 861.5
}
Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry".
Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns 422 with a message.
Generation: POST /v1/chat/completions
The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes messages (OpenAI-style roles) and max_tokens. Decoding is greedy.
{"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40}
{
"object": "chat.completion",
"choices": [{"index": 0, "finish_reason": "stop",
"message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}],
"tokens": [32, 20965, 374, "..."],
"usage": {"completion_tokens": 30}
}
tokens carries the raw token ids, which is how the identity checks compare outputs.
Decide, then generate if needed: POST /v1/auto
One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (state_prefills is always 1).
{
"state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.",
"name": "count",
"question": {"type": "choice", "instructions": "How many stamps does Tara have now?",
"criteria": {"12": null, "13": null, "24": null, "14": null}},
"max_tokens": 300
}
{
"decision": {"type": "choice", "choice": "14",
"probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587},
"router_score": 0.855, "needs_generation": true},
"escalated": true,
"state_prefills": 1,
"generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A",
"generation_tokens": ["126 ids"]
}
Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When escalated is false the response has only decision.
Speed
| First token after escalation, cache reuse | Encoding the document again | |
|---|---|---|
| Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s |
| Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s |
| H100, bf16, 4,091-token document | 33 ms | 98 ms |
| H100, bf16, 257-token document | 32 ms | 29 ms |
A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100.
Limits
- One request on the model at a time; requests queue.
- English, and documents up to about 4,000 characters were trained on.
- Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what
autois for.