Dev-4B / docs /api.md
suhaas-teja's picture
Dev-4B v0.1.0
a69969c verified
|
Raw History Blame Contribute Delete
6.13 kB
# API
One server process answers three kinds of request. The GPU server (`serve/server.py`, PyTorch) and the Mac server (`serve/server_mlx.py`, MLX) share the same routes and formats.
| Route | What it does |
| --- | --- |
| `POST /v1/systemone` | Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag |
| `POST /v1/chat/completions` | Plain chat with the base model (the adapter is off) |
| `POST /v1/auto` | One question: decide, and if the router flags it, answer by generating from the same document cache |
Start a server:
```bash
# From the downloaded Dev-4B folder (see README.md)
python -m serve.server --artifacts . # NVIDIA GPU
python -m serve.server_mlx --model mlx-8bit --artifacts . # Mac (Dev-4B-MLX-8bit in mlx-8bit)
```
The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly.
## Decisions: `POST /v1/systemone`
The *state* is the document. Each question is one of three types:
| Type | Options | Answer fields |
| --- | --- | --- |
| `noul` | yes / no | `noul`: probability of yes |
| `choice` | `criteria`: 1 to 16 named options, each with an optional description | `choice`: the most likely option name |
| `score` | `criteria`: 2 to 6 ordered level descriptions, lowest first | `score`: expected level, 0 = first |
Every answer also has `probabilities`, a `router_score` (0 to 1, higher means a single pass is more likely to be wrong) and `needs_generation` (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state.
Request:
```json
{
"state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
"tone": {"type": "score", "instructions": "How upset is the customer?",
"criteria": ["calm", "annoyed", "angry"]}
}
}
```
Response (trimmed):
```json
{
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009},
"router_score": 0.246, "needs_generation": false},
"urgent": {"type": "noul", "noul": 0.615,
"probabilities": {"false": 0.385, "true": 0.615},
"router_score": 0.799, "needs_generation": true},
"tone": {"type": "score", "score": 1.46,
"probabilities": {"0": 0.038, "1": 0.462, "2": 0.500},
"router_score": 0.363, "needs_generation": false}
},
"latency_ms": 861.5
}
```
Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry".
Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns `422` with a message.
## Generation: `POST /v1/chat/completions`
The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes `messages` (OpenAI-style roles) and `max_tokens`. Decoding is greedy.
```json
{"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40}
```
```json
{
"object": "chat.completion",
"choices": [{"index": 0, "finish_reason": "stop",
"message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}],
"tokens": [32, 20965, 374, "..."],
"usage": {"completion_tokens": 30}
}
```
`tokens` carries the raw token ids, which is how the identity checks compare outputs.
## Decide, then generate if needed: `POST /v1/auto`
One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (`state_prefills` is always 1).
```json
{
"state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.",
"name": "count",
"question": {"type": "choice", "instructions": "How many stamps does Tara have now?",
"criteria": {"12": null, "13": null, "24": null, "14": null}},
"max_tokens": 300
}
```
```json
{
"decision": {"type": "choice", "choice": "14",
"probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587},
"router_score": 0.855, "needs_generation": true},
"escalated": true,
"state_prefills": 1,
"generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A",
"generation_tokens": ["126 ids"]
}
```
Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When `escalated` is false the response has only `decision`.
## Speed
| | First token after escalation, cache reuse | Encoding the document again |
| --- | --- | --- |
| Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s |
| Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s |
| H100, bf16, 4,091-token document | 33 ms | 98 ms |
| H100, bf16, 257-token document | 32 ms | 29 ms |
A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100.
## Limits
- One request on the model at a time; requests queue.
- English, and documents up to about 4,000 characters were trained on.
- Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what `auto` is for.