decision-studio / SYSTEMONE_API.md
Xunzhuo's picture
Cursor
Add Decision 2.0 models to the Gateway registry
e9ab6b6
|
Raw History Blame Contribute Delete
9.1 kB

Decision SystemOne API

Call the Decision 2.0 models (Vega, Lux, Nox, Sol, Eos and Kai) and the Decision 1.0 models (Lux, Nox, Sol, Eos, Kai and Lex) from your computer using the same state / questions / model request format as the SystemOne API.

  • Base URL: https://YOUR-DECISION-HOST (replace with your deployment origin)
  • Single-state inference: POST /v1/systemone
  • Shared-question batches: POST /v1/systemone/batches (Decision extension)
  • Model discovery: GET /v1/models

This is a public demo endpoint backed by the released models on AMD GPUs. It requires no API key. Inference runs on the connected server; your computer sends the request. It is a shared, bounded service rather than a dedicated production deployment. Responses are real model outputs; there is no chat-completion endpoint.

Send a request

curl --fail-with-body https://YOUR-DECISION-HOST/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"model":"vllm-sr/Decision-2.0-Lux-9B","state":"My subscription was charged twice. Please refund the duplicate charge.","questions":{"billing":{"type":"noul","instructions":"Does the customer report a billing problem?"},"team":{"type":"choice","instructions":"Which team should handle this message?","criteria":{"Billing":"Charges, invoices, and refunds","Accounts":"Login, passwords, and account access"}},"urgency":{"type":"score","instructions":"Rate how urgently this needs attention.","criteria":["Routine","Needs prompt attention","Critical"]}}}'

Ask multiple named questions in one call. The same state is used for every question and the worker processes the admitted decisions in GPU batches. This does not imply shared-state encoder caching.

/v1/systemone/batches keeps the same model, question types, and per-state answer shapes as /v1/systemone, but replaces the single state with an ordered states: [{id, state}, ...] array. It returns ordered results instead of one top-level answers map. This is our multi-state extension; the official SystemOne SDKs call only the single-state endpoint.

To apply one question set to multiple independent states, send a batch:

curl --fail-with-body https://YOUR-DECISION-HOST/v1/systemone/batches \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "vllm-sr/Decision-2.0-Lux-9B",
    "states": [
      {"id": "ticket-a", "state": "I was charged twice. Please refund the duplicate."},
      {"id": "ticket-b", "state": "How do I change my password?"}
    ],
    "questions": {
      "billing": {"type": "noul", "instructions": "Is this a billing issue?"},
      "team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"Billing": "Charges and refunds", "Accounts": "Account access"}}
    }
  }'

Read results[0].id and results[0].answers for ticket-a, then results[1] for ticket-b. The results array preserves input order and each row has its own usage; top-level usage is the total. This is one request with four decisions, not the same as four HTTP requests.

Official Python SDK

Use Python 3.10 or later:

python -m pip install typesafe-sdk==0.7.1
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient, RetryPolicy

with TypeSafeClient(
    base_url="https://YOUR-DECISION-HOST",
    api_key="decision-public",  # SDK-required placeholder, not a credential
    model="vllm-sr/Decision-2.0-Lux-9B",
    timeout=60,
    retry=RetryPolicy(max_retries=0),
) as client:
    print([model.name for model in client.models.list().models])
    result = client.system_one(
        state="My subscription was charged twice. Please refund the duplicate charge.",
        questions={
            "billing": Noul(instructions="Does the customer report a billing problem?"),
            "team": Choice(
                instructions="Which team should handle this message?",
                criteria={"Billing": "Charges, invoices, and refunds",
                          "Accounts": "Login, passwords, and account access"},
            ),
            "urgency": Score(
                instructions="Rate how urgently this needs attention.",
                criteria=["Routine", "Needs prompt attention", "Critical"],
            ),
        },
    )
    print(result.nouls["billing"].noul)
    print(result.choices["team"].choice)
    print(result.scores["urgency"].score)

Select any exact Hugging Face repository ID returned as id by /v1/models. The model is required on public inference requests; short names, wire aliases, and case variants are rejected. The response repeats the selected canonical ID. The Studio editor uses a same-origin endpoint; in direct serving mode it receives the same strict response envelope, while legacy modes retain their internal queue/diagnostic contracts.

Direct mode does not advertise a default model. The Studio editor selects a model locally and includes it explicitly in every /api/evaluate request; a missing model receives HTTP 422. Legacy rollback modes may retain their previous discovery default.

Model Generation Parameters Complete input per question
Vega 2.0 27B 32,768 tokens
Lux 2.0 9B 16,384 tokens
Nox 2.0 4B 16,384 tokens
Sol 2.0 2B 16,384 tokens
Eos 2.0 0.8B 16,384 tokens
Kai 2.0 0.6B 8,192 tokens
Lux 1.0 9B 16,384 tokens
Nox 1.0 4B 16,384 tokens
Sol 1.0 2B 16,384 tokens
Eos 1.0 0.8B 16,384 tokens
Kai 1.0 0.6B 1,024 tokens
Lex 1.0 0.6B 1,024 tokens

GET /v1/models returns models in this order, including pinned revision/manifest metadata. When present, release_date comes from MODEL_RELEASES.json and applies to the currently configured Hub revision and manifest. The date is omitted for a newly deployed artifact until its presentation metadata is updated. If that presentation file is missing or invalid, the endpoint still lists all configured models without release dates; serving does not depend on the file. Explicit metadata validation still reports the error.

Responses and limits

  • Noul returns P(true); Choice returns its winning label and probabilities; Score returns the expected ordinal level and its distribution.
  • confidence is a versioned statistic of the answer distribution, named per model by confidence_definition in /v1/models. Decision 1.0 (margin_ordinal_v1): Choice reports the top-two probability margin and Score the ordinal concentration around its expected level, normalized against uniform variance. Decision 2.0 (normalized_entropy_v2): Choice and Score report one minus the distribution's entropy divided by the log of its candidate count. These are Decision-owned statistics, not calibrated probabilities of correctness or a claim of equivalence to TypeSafe's undisclosed calculation. The direct Gateway validates and forwards the runtime value unchanged; it never substitutes max(probabilities). Public responses contain only model, answers, and usage.
  • The hosted endpoint accepts one or more named questions without a fixed question-count cap. In direct mode, each model instance owns its physical microbatch setting. Choice accepts 2–255 options, Score accepts 2–10 ordered levels, and JSON requests are limited to 256 KiB. Expanded state/question input is limited to 16 MiB so large Cartesian workloads stay bounded. All question types require nonempty instructions. State plus question plus all candidates must fit the model limit; overflowing requests are rejected without truncation.
  • In direct serving mode, the synchronous wait is configurable and each instance owns its own concurrency, queue, and physical-batch limits; /v1/models does not invent fixed values for them. A busy or timed-out service may return 529/503/504 (legacy queue mode may return 429). Completed requests do not consume admission capacity. Do not treat transport errors as predictions. SDK example retries are disabled to avoid duplicate inference.
  • POST /v1/systemone/batches applies one shared questions map to ordered states: [{id, state}, ...]. It requires the same explicit canonical model, unique state IDs, at most 1,024 states, questions, and total decisions, a 2 MiB JSON body, and the same 16 MiB expanded-input and per-question token budgets. Its atomic response contains model, ordered results: [{id, answers, usage}, ...], and aggregate usage. This Decision extension is separate from the SDK's single-state method.

See SYSTEM_ONE_MAPPING.md for the complete adapter contract. The standard single-state call, all three answer types, and model discovery are validated with the official SDK.

In legacy pull-queue mode, completed asynchronous results are kept for up to 120 seconds, with the oldest completed records evicted when the 64-result cache fills. Direct serving mode does not create Gateway jobs or retain completion receipts. GET /v1/models publishes the service limits. The original Space origin remains available.