Text Classification
MLX
English
decision-model
calibration
lora
activated-lora
adaptive-reasoning
on-device
Instructions to use suhaas-teja/Dev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use suhaas-teja/Dev-4B with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download suhaas-teja/Dev-4B --local-dir Dev-4B
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
|
Download docs/api.md from suhaas-teja/Dev-4B: direct link, hf CLI and curl.
- Browser
- Download file 6.13 kB
-
https://huggingface.co/suhaas-teja/Dev-4B/resolve/main/docs/api.md
- Command line
-
hf download hf://suhaas-teja/Dev-4B/docs/api.md
-
curl -L -o api.md https://huggingface.co/suhaas-teja/Dev-4B/resolve/main/docs/api.md
6.13 kB
| # API | |
| One server process answers three kinds of request. The GPU server (`serve/server.py`, PyTorch) and the Mac server (`serve/server_mlx.py`, MLX) share the same routes and formats. | |
| | Route | What it does | | |
| | --- | --- | | |
| | `POST /v1/systemone` | Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag | | |
| | `POST /v1/chat/completions` | Plain chat with the base model (the adapter is off) | | |
| | `POST /v1/auto` | One question: decide, and if the router flags it, answer by generating from the same document cache | | |
| Start a server: | |
| ```bash | |
| # From the downloaded Dev-4B folder (see README.md) | |
| python -m serve.server --artifacts . # NVIDIA GPU | |
| python -m serve.server_mlx --model mlx-8bit --artifacts . # Mac (Dev-4B-MLX-8bit in mlx-8bit) | |
| ``` | |
| The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly. | |
| ## Decisions: `POST /v1/systemone` | |
| The *state* is the document. Each question is one of three types: | |
| | Type | Options | Answer fields | | |
| | --- | --- | --- | | |
| | `noul` | yes / no | `noul`: probability of yes | | |
| | `choice` | `criteria`: 1 to 16 named options, each with an optional description | `choice`: the most likely option name | | |
| | `score` | `criteria`: 2 to 6 ordered level descriptions, lowest first | `score`: expected level, 0 = first | | |
| Every answer also has `probabilities`, a `router_score` (0 to 1, higher means a single pass is more likely to be wrong) and `needs_generation` (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state. | |
| Request: | |
| ```json | |
| { | |
| "state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.", | |
| "questions": { | |
| "team": {"type": "choice", "instructions": "Which team should handle this?", | |
| "criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}}, | |
| "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}, | |
| "tone": {"type": "score", "instructions": "How upset is the customer?", | |
| "criteria": ["calm", "annoyed", "angry"]} | |
| } | |
| } | |
| ``` | |
| Response (trimmed): | |
| ```json | |
| { | |
| "answers": { | |
| "team": {"type": "choice", "choice": "billing", | |
| "probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009}, | |
| "router_score": 0.246, "needs_generation": false}, | |
| "urgent": {"type": "noul", "noul": 0.615, | |
| "probabilities": {"false": 0.385, "true": 0.615}, | |
| "router_score": 0.799, "needs_generation": true}, | |
| "tone": {"type": "score", "score": 1.46, | |
| "probabilities": {"0": 0.038, "1": 0.462, "2": 0.500}, | |
| "router_score": 0.363, "needs_generation": false} | |
| }, | |
| "latency_ms": 861.5 | |
| } | |
| ``` | |
| Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry". | |
| Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns `422` with a message. | |
| ## Generation: `POST /v1/chat/completions` | |
| The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes `messages` (OpenAI-style roles) and `max_tokens`. Decoding is greedy. | |
| ```json | |
| {"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40} | |
| ``` | |
| ```json | |
| { | |
| "object": "chat.completion", | |
| "choices": [{"index": 0, "finish_reason": "stop", | |
| "message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}], | |
| "tokens": [32, 20965, 374, "..."], | |
| "usage": {"completion_tokens": 30} | |
| } | |
| ``` | |
| `tokens` carries the raw token ids, which is how the identity checks compare outputs. | |
| ## Decide, then generate if needed: `POST /v1/auto` | |
| One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (`state_prefills` is always 1). | |
| ```json | |
| { | |
| "state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.", | |
| "name": "count", | |
| "question": {"type": "choice", "instructions": "How many stamps does Tara have now?", | |
| "criteria": {"12": null, "13": null, "24": null, "14": null}}, | |
| "max_tokens": 300 | |
| } | |
| ``` | |
| ```json | |
| { | |
| "decision": {"type": "choice", "choice": "14", | |
| "probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587}, | |
| "router_score": 0.855, "needs_generation": true}, | |
| "escalated": true, | |
| "state_prefills": 1, | |
| "generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A", | |
| "generation_tokens": ["126 ids"] | |
| } | |
| ``` | |
| Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When `escalated` is false the response has only `decision`. | |
| ## Speed | |
| | | First token after escalation, cache reuse | Encoding the document again | | |
| | --- | --- | --- | | |
| | Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s | | |
| | Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s | | |
| | H100, bf16, 4,091-token document | 33 ms | 98 ms | | |
| | H100, bf16, 257-token document | 32 ms | 29 ms | | |
| A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100. | |
| ## Limits | |
| - One request on the model at a time; requests queue. | |
| - English, and documents up to about 4,000 characters were trained on. | |
| - Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what `auto` is for. | |