# Decision SystemOne API Call the Decision 2.0 models (Vega, Lux, Nox, Sol, Eos and Kai) and the Decision 1.0 models (Lux, Nox, Sol, Eos, Kai and Lex) from your computer using the same `state / questions / model` request format as the [SystemOne API](https://docs.typesafe.ai/api). - Base URL: `https://YOUR-DECISION-HOST` (replace with your deployment origin) - Single-state inference: `POST /v1/systemone` - Shared-question batches: `POST /v1/systemone/batches` (Decision extension) - Model discovery: `GET /v1/models` This is a public demo endpoint backed by the released models on AMD GPUs. It requires no API key. Inference runs on the connected server; your computer sends the request. It is a shared, bounded service rather than a dedicated production deployment. Responses are real model outputs; there is no chat-completion endpoint. ## Send a request ```bash curl --fail-with-body https://YOUR-DECISION-HOST/v1/systemone \ -H 'Content-Type: application/json' \ -d '{"model":"vllm-sr/Decision-2.0-Lux-9B","state":"My subscription was charged twice. Please refund the duplicate charge.","questions":{"billing":{"type":"noul","instructions":"Does the customer report a billing problem?"},"team":{"type":"choice","instructions":"Which team should handle this message?","criteria":{"Billing":"Charges, invoices, and refunds","Accounts":"Login, passwords, and account access"}},"urgency":{"type":"score","instructions":"Rate how urgently this needs attention.","criteria":["Routine","Needs prompt attention","Critical"]}}}' ``` Ask multiple named questions in one call. The same state is used for every question and the worker processes the admitted decisions in GPU batches. This does not imply shared-state encoder caching. `/v1/systemone/batches` keeps the same `model`, question types, and per-state answer shapes as `/v1/systemone`, but replaces the single `state` with an ordered `states: [{id, state}, ...]` array. It returns ordered `results` instead of one top-level `answers` map. This is our multi-state extension; the official SystemOne SDKs call only the single-state endpoint. To apply one question set to multiple independent states, send a batch: ```bash curl --fail-with-body https://YOUR-DECISION-HOST/v1/systemone/batches \ -H 'Content-Type: application/json' \ -d '{ "model": "vllm-sr/Decision-2.0-Lux-9B", "states": [ {"id": "ticket-a", "state": "I was charged twice. Please refund the duplicate."}, {"id": "ticket-b", "state": "How do I change my password?"} ], "questions": { "billing": {"type": "noul", "instructions": "Is this a billing issue?"}, "team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"Billing": "Charges and refunds", "Accounts": "Account access"}} } }' ``` Read `results[0].id` and `results[0].answers` for `ticket-a`, then `results[1]` for `ticket-b`. The `results` array preserves input order and each row has its own `usage`; top-level `usage` is the total. This is one request with four decisions, not the same as four HTTP requests. ## Official Python SDK Use Python 3.10 or later: ```bash python -m pip install typesafe-sdk==0.7.1 ``` ```python from typesafe_sdk import Choice, Noul, Score, TypeSafeClient, RetryPolicy with TypeSafeClient( base_url="https://YOUR-DECISION-HOST", api_key="decision-public", # SDK-required placeholder, not a credential model="vllm-sr/Decision-2.0-Lux-9B", timeout=60, retry=RetryPolicy(max_retries=0), ) as client: print([model.name for model in client.models.list().models]) result = client.system_one( state="My subscription was charged twice. Please refund the duplicate charge.", questions={ "billing": Noul(instructions="Does the customer report a billing problem?"), "team": Choice( instructions="Which team should handle this message?", criteria={"Billing": "Charges, invoices, and refunds", "Accounts": "Login, passwords, and account access"}, ), "urgency": Score( instructions="Rate how urgently this needs attention.", criteria=["Routine", "Needs prompt attention", "Critical"], ), }, ) print(result.nouls["billing"].noul) print(result.choices["team"].choice) print(result.scores["urgency"].score) ``` Select any exact Hugging Face repository ID returned as `id` by `/v1/models`. The model is required on public inference requests; short names, wire aliases, and case variants are rejected. The response repeats the selected canonical ID. The Studio editor uses a same-origin endpoint; in direct serving mode it receives the same strict response envelope, while legacy modes retain their internal queue/diagnostic contracts. Direct mode does not advertise a default model. The Studio editor selects a model locally and includes it explicitly in every `/api/evaluate` request; a missing model receives HTTP 422. Legacy rollback modes may retain their previous discovery default. | Model | Generation | Parameters | Complete input per question | |---|---|---:|---:| | Vega | 2.0 | 27B | 32,768 tokens | | Lux | 2.0 | 9B | 16,384 tokens | | Nox | 2.0 | 4B | 16,384 tokens | | Sol | 2.0 | 2B | 16,384 tokens | | Eos | 2.0 | 0.8B | 16,384 tokens | | Kai | 2.0 | 0.6B | 8,192 tokens | | Lux | 1.0 | 9B | 16,384 tokens | | Nox | 1.0 | 4B | 16,384 tokens | | Sol | 1.0 | 2B | 16,384 tokens | | Eos | 1.0 | 0.8B | 16,384 tokens | | Kai | 1.0 | 0.6B | 1,024 tokens | | Lex | 1.0 | 0.6B | 1,024 tokens | `GET /v1/models` returns models in this order, including pinned revision/manifest metadata. When present, `release_date` comes from `MODEL_RELEASES.json` and applies to the currently configured Hub revision and manifest. The date is omitted for a newly deployed artifact until its presentation metadata is updated. If that presentation file is missing or invalid, the endpoint still lists all configured models without release dates; serving does not depend on the file. Explicit metadata validation still reports the error. ## Responses and limits - Noul returns P(true); Choice returns its winning label and probabilities; Score returns the expected ordinal level and its distribution. - `confidence` is a versioned statistic of the answer distribution, named per model by `confidence_definition` in `/v1/models`. Decision 1.0 (`margin_ordinal_v1`): Choice reports the top-two probability margin and Score the ordinal concentration around its expected level, normalized against uniform variance. Decision 2.0 (`normalized_entropy_v2`): Choice and Score report one minus the distribution's entropy divided by the log of its candidate count. These are Decision-owned statistics, not calibrated probabilities of correctness or a claim of equivalence to TypeSafe's undisclosed calculation. The direct Gateway validates and forwards the runtime value unchanged; it never substitutes `max(probabilities)`. Public responses contain only `model`, `answers`, and `usage`. - The hosted endpoint accepts one or more named questions without a fixed question-count cap. In direct mode, each model instance owns its physical microbatch setting. Choice accepts 2–255 options, Score accepts 2–10 ordered levels, and JSON requests are limited to 256 KiB. Expanded state/question input is limited to 16 MiB so large Cartesian workloads stay bounded. All question types require nonempty instructions. State plus question plus all candidates must fit the model limit; overflowing requests are rejected without truncation. - In direct serving mode, the synchronous wait is configurable and each instance owns its own concurrency, queue, and physical-batch limits; `/v1/models` does not invent fixed values for them. A busy or timed-out service may return 529/503/504 (legacy queue mode may return 429). Completed requests do not consume admission capacity. Do not treat transport errors as predictions. SDK example retries are disabled to avoid duplicate inference. - `POST /v1/systemone/batches` applies one shared `questions` map to ordered `states: [{id, state}, ...]`. It requires the same explicit canonical `model`, unique state IDs, at most 1,024 states, questions, and total decisions, a 2 MiB JSON body, and the same 16 MiB expanded-input and per-question token budgets. Its atomic response contains `model`, ordered `results: [{id, answers, usage}, ...]`, and aggregate `usage`. This Decision extension is separate from the SDK's single-state method. See [SYSTEM_ONE_MAPPING.md](SYSTEM_ONE_MAPPING.md) for the complete adapter contract. The standard single-state call, all three answer types, and model discovery are validated with the official SDK. In legacy pull-queue mode, completed asynchronous results are kept for up to 120 seconds, with the oldest completed records evicted when the 64-result cache fills. Direct serving mode does not create Gateway jobs or retain completion receipts. `GET /v1/models` publishes the service limits. The original Space origin remains available.