Dev-4B / docs /api.md
suhaas-teja's picture
Dev-4B v0.1.0
a69969c verified
|
Raw History Blame Contribute Delete
6.13 kB

API

One server process answers three kinds of request. The GPU server (serve/server.py, PyTorch) and the Mac server (serve/server_mlx.py, MLX) share the same routes and formats.

Route What it does
POST /v1/systemone Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag
POST /v1/chat/completions Plain chat with the base model (the adapter is off)
POST /v1/auto One question: decide, and if the router flags it, answer by generating from the same document cache

Start a server:

# From the downloaded Dev-4B folder (see README.md)
python -m serve.server --artifacts .                       # NVIDIA GPU
python -m serve.server_mlx --model mlx-8bit --artifacts .  # Mac (Dev-4B-MLX-8bit in mlx-8bit)

The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly.

Decisions: POST /v1/systemone

The state is the document. Each question is one of three types:

Type Options Answer fields
noul yes / no noul: probability of yes
choice criteria: 1 to 16 named options, each with an optional description choice: the most likely option name
score criteria: 2 to 6 ordered level descriptions, lowest first score: expected level, 0 = first

Every answer also has probabilities, a router_score (0 to 1, higher means a single pass is more likely to be wrong) and needs_generation (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state.

Request:

{
  "state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.",
  "questions": {
    "team":   {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}},
    "urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
    "tone":   {"type": "score", "instructions": "How upset is the customer?",
               "criteria": ["calm", "annoyed", "angry"]}
  }
}

Response (trimmed):

{
  "answers": {
    "team":   {"type": "choice", "choice": "billing",
               "probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009},
               "router_score": 0.246, "needs_generation": false},
    "urgent": {"type": "noul", "noul": 0.615,
               "probabilities": {"false": 0.385, "true": 0.615},
               "router_score": 0.799, "needs_generation": true},
    "tone":   {"type": "score", "score": 1.46,
               "probabilities": {"0": 0.038, "1": 0.462, "2": 0.500},
               "router_score": 0.363, "needs_generation": false}
  },
  "latency_ms": 861.5
}

Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry".

Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns 422 with a message.

Generation: POST /v1/chat/completions

The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes messages (OpenAI-style roles) and max_tokens. Decoding is greedy.

{"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40}
{
  "object": "chat.completion",
  "choices": [{"index": 0, "finish_reason": "stop",
               "message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}],
  "tokens": [32, 20965, 374, "..."],
  "usage": {"completion_tokens": 30}
}

tokens carries the raw token ids, which is how the identity checks compare outputs.

Decide, then generate if needed: POST /v1/auto

One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (state_prefills is always 1).

{
  "state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.",
  "name": "count",
  "question": {"type": "choice", "instructions": "How many stamps does Tara have now?",
               "criteria": {"12": null, "13": null, "24": null, "14": null}},
  "max_tokens": 300
}
{
  "decision": {"type": "choice", "choice": "14",
               "probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587},
               "router_score": 0.855, "needs_generation": true},
  "escalated": true,
  "state_prefills": 1,
  "generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A",
  "generation_tokens": ["126 ids"]
}

Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When escalated is false the response has only decision.

Speed

First token after escalation, cache reuse Encoding the document again
Mac, M1 Pro, 8-bit, 1,013-token document 0.30 s 3.4 s
Mac, M1 Pro, 8-bit, 4,091-token document 0.39 s 13.9 s
H100, bf16, 4,091-token document 33 ms 98 ms
H100, bf16, 257-token document 32 ms 29 ms

A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100.

Limits

  • One request on the model at a time; requests queue.
  • English, and documents up to about 4,000 characters were trained on.
  • Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what auto is for.