File size: 6,128 Bytes
a69969c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
# API

One server process answers three kinds of request. The GPU server (`serve/server.py`, PyTorch) and the Mac server (`serve/server_mlx.py`, MLX) share the same routes and formats.

| Route | What it does |
| --- | --- |
| `POST /v1/systemone` | Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag |
| `POST /v1/chat/completions` | Plain chat with the base model (the adapter is off) |
| `POST /v1/auto` | One question: decide, and if the router flags it, answer by generating from the same document cache |

Start a server:

```bash
# From the downloaded Dev-4B folder (see README.md)
python -m serve.server --artifacts .                       # NVIDIA GPU
python -m serve.server_mlx --model mlx-8bit --artifacts .  # Mac (Dev-4B-MLX-8bit in mlx-8bit)
```

The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly.

## Decisions: `POST /v1/systemone`

The *state* is the document. Each question is one of three types:

| Type | Options | Answer fields |
| --- | --- | --- |
| `noul` | yes / no | `noul`: probability of yes |
| `choice` | `criteria`: 1 to 16 named options, each with an optional description | `choice`: the most likely option name |
| `score` | `criteria`: 2 to 6 ordered level descriptions, lowest first | `score`: expected level, 0 = first |

Every answer also has `probabilities`, a `router_score` (0 to 1, higher means a single pass is more likely to be wrong) and `needs_generation` (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state.

Request:

```json
{
  "state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.",
  "questions": {
    "team":   {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}},
    "urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
    "tone":   {"type": "score", "instructions": "How upset is the customer?",
               "criteria": ["calm", "annoyed", "angry"]}
  }
}
```

Response (trimmed):

```json
{
  "answers": {
    "team":   {"type": "choice", "choice": "billing",
               "probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009},
               "router_score": 0.246, "needs_generation": false},
    "urgent": {"type": "noul", "noul": 0.615,
               "probabilities": {"false": 0.385, "true": 0.615},
               "router_score": 0.799, "needs_generation": true},
    "tone":   {"type": "score", "score": 1.46,
               "probabilities": {"0": 0.038, "1": 0.462, "2": 0.500},
               "router_score": 0.363, "needs_generation": false}
  },
  "latency_ms": 861.5
}
```

Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry".

Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns `422` with a message.

## Generation: `POST /v1/chat/completions`

The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes `messages` (OpenAI-style roles) and `max_tokens`. Decoding is greedy.

```json
{"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40}
```

```json
{
  "object": "chat.completion",
  "choices": [{"index": 0, "finish_reason": "stop",
               "message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}],
  "tokens": [32, 20965, 374, "..."],
  "usage": {"completion_tokens": 30}
}
```

`tokens` carries the raw token ids, which is how the identity checks compare outputs.

## Decide, then generate if needed: `POST /v1/auto`

One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (`state_prefills` is always 1).

```json
{
  "state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.",
  "name": "count",
  "question": {"type": "choice", "instructions": "How many stamps does Tara have now?",
               "criteria": {"12": null, "13": null, "24": null, "14": null}},
  "max_tokens": 300
}
```

```json
{
  "decision": {"type": "choice", "choice": "14",
               "probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587},
               "router_score": 0.855, "needs_generation": true},
  "escalated": true,
  "state_prefills": 1,
  "generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A",
  "generation_tokens": ["126 ids"]
}
```

Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When `escalated` is false the response has only `decision`.

## Speed

| | First token after escalation, cache reuse | Encoding the document again |
| --- | --- | --- |
| Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s |
| Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s |
| H100, bf16, 4,091-token document | 33 ms | 98 ms |
| H100, bf16, 257-token document | 32 ms | 29 ms |

A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100.

## Limits

- One request on the model at a time; requests queue.
- English, and documents up to about 4,000 characters were trained on.
- Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what `auto` is for.