sys1-micro

68 million parameters for fast System One decisions: classify, route, score, and answer multiple questions over one state. Returns typed probabilities for Choice, Score, and Noul (yes/no).

Trained on 2,520,976 states / 3,739,677 questions. Indexed test states [question/prompt text] and constituents [source passages within composite states] were excluded from the accepted training build.

Install and run

pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sulabhkatiyar/sys1-micro', local_dir='./sys1-micro')"
pip install ./sys1-micro
from sysone import SysOne

model = SysOne.from_pretrained("./sys1-micro")
answers = model.predict([{
    "state": "A customer was charged twice for the same purchase and requests a refund.",
    "questions": [
        {"type": "choice", "instructions": "Which team should handle this request?",
         "options": ["billing", "technical support", "sales"]},
        {"type": "noul", "instructions": "Does the customer report a duplicate charge?"}
    ]
}])
print(answers)

Tested answers for this model: billing (probability 0.9458) and yes (probability 0.9490).

For GPU inference on a CUDA-enabled machine:

model = SysOne.from_pretrained("./sys1-micro", device="cuda")

Both models were tested on one Tesla T4 in fp32 with the default 2,048-token context.

Loading requires the custom SysOne client; Transformers AutoModel loading is unsupported. config.json configures the native encoder and decision heads, and architectures: ["SysOne"] identifies this package's class.

Choice/Score return probs in input option order and argmax; Noul returns p for true. Score also returns the expected zero-based level index as value: sum(probs[k] * k). State may be text or structured JSON. Defaults: CPU fp32, four threads, 2,048-token runtime context. The optional context_length=4096 executes the full 231-item public JevBench panel without truncation; training context remains 2,048. Over-length questions return error: too_long.

Benchmarks

LangWatch mean of eleven native task metrics

JevBench public aggregate

BFCL selection comparison

JevBench size and accuracy

The LangWatch summary is the mean of eleven native task metrics; panels use source/count-inferred samples. JevBench is full public231 raw accuracy at the declared 4,096 runtime guard. BFCL is function selection, with its stated variant scoring. Peer values are sourced reports. Selected tested requests and actual probabilities are in release_examples.json.

Appendix: task panels

LangWatch eleven task metrics

JevBench families 1

JevBench families 2

JevBench families 3

License

MIT License. See LICENSE for the full text and preserved upstream notices.

Citation

@misc{katiyar2026sys1micro,
  author = {Sulabh Katiyar},
  title = {sys1 micro},
  year = {2026},
  url = {https://huggingface.co/sulabhkatiyar/sys1-micro}
}
Downloads last month
6
Safetensors
Model size
68.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sulabhkatiyar/sys1-micro

Finetuned
(23)
this model

Collection including sulabhkatiyar/sys1-micro