Instructions to use newcombs/onebox-1-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use newcombs/onebox-1-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="newcombs/onebox-1-9b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("newcombs/onebox-1-9b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Onebox 1 (9B)
Onebox makes typed decisions about a state. You send a state (text or JSON) and one or more questions, each with a fixed set of allowed options. You get back a probability for every option of every question. The model never writes free text, so there is nothing to parse and nothing to repair.
The name comes from Newcomb's problem, where the decision is whether to take one box or two. Onebox takes its decision in one pass.
This is the first release of the Onebox series from the newcombs lab. It is a research release: we use it to study how well a local model makes typed decisions and how often it falls for misleading words in the input. This is our measure of whether a model has learned to read a text closely and understand it, rather than to just pick the option whose word fits best.
What is in this release
onebox.safetensors: weights we trained for Qwen3.5-9B, 44 million parameters (low rank adapters on the last 18 layers and a small scoring head). Only the text part of the base model is used.onebox.py: the inference code. One file, no training code.onebox_config.json: base model, sizes and the calibration temperature.
Question types
| type | what you send | what you get |
|---|---|---|
choice |
named options, optionally with a description each | a probability per option |
noul |
a yes or no question | the probability of yes |
score |
ordered levels | a probability per level and the expected level |
The order in which you list the options does not change the answer. Every option is read on its own against the same state, and the options are compared only at the end.
How to run it
pip install torch "transformers>=5.10" safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("newcombs/onebox-1-9b")
sys.path.insert(0, path)
from onebox import load
model = load(path) # fetches Qwen/Qwen3.5-9B on first use, about 19 GB
print(model.decide(
{"ticket": "Server down since 3am, customers cannot log in"},
{
"team": {
"type": "choice",
"instructions": "Which team handles this?",
"criteria": {"billing": "payments, invoices", "technical": "outages, bugs, login", "sales": "new contracts"},
},
"urgent": {"type": "noul", "instructions": "Does this need action today?"},
"severity": {"type": "score", "instructions": "How severe is this?", "criteria": ["minor", "major", "critical"]},
},
))
The request and the answer have the same shape as TypeSafe's System One API (state, questions with type, instructions and criteria). load picks CUDA, Apple MPS or the CPU on its own. In bf16 the model needs about 18 GB of memory.
Output (rounded to four places):
{'answers': {
'team': {'type': 'choice', 'choice': 'technical', 'confidence': 0.9307,
'probabilities': {'billing': 0.035, 'technical': 0.9307, 'sales': 0.0343}},
'urgent': {'type': 'noul', 'noul': 0.9774},
'severity': {'type': 'score', 'score': 1.3782, 'confidence': 0.5991,
'legend': {'0': 'minor', '1': 'major', '2': 'critical'},
'probabilities': {'0': 0.0114, '1': 0.5991, '2': 0.3895}}}}
Results
All numbers are measured by us with one harness. A decision counts as correct when the most likely option matches the label.
typed-decisions
Public benchmark LocalLLaMA/typed-decisions, official test split: 400 cases with 2000 decisions from four workflows (agent traces, customer service, invoices, security incidents).
| model | trained on these four workflows? | accuracy |
|---|---|---|
| always the most common answer | n/a | 46.1 % |
| Laya 421M | no | 36.2 % |
| Onebox recipe without the typed-decisions training split | no | 59.4 % |
| openjev 4B v5 | no, according to its manifest | 63.8 % |
| Clef-Flash (9B) | not as far as its card says | 70.3 % |
| Julia-1 | unknown | 72.6 % |
| TypeSafe Jev (closed, published number) | unknown | 72.7 % |
| Laya 421M, fine-tuned | yes | 76.6 % |
| Onebox 1 (9B) | yes | 79.3 % |
Onebox was trained on the training split of this benchmark, so the 79.3 % is the number for workflows it has seen. The row without the training split is closer to what to expect on a workflow it has never seen.
By question type on the test split, Onebox reaches 75.3 % on choice, 86.2 % on yes or no questions and 77.1 % on score. Its expected calibration error is 0.150. On the validation split it reaches 80.5 %.
How this model was chosen
We trained seven variants with the typed-decisions training split and picked the one with the best accuracy on the typed-decisions validation split, among the variants that fell for at most 10 of our 104 older lure cases. The released model was trained twice as long as our first variant (7200 instead of 3600 steps).
Misleading words
A lure is a word in the input that points toward the wrong option. A customer writes "I can't open the billing page." The page does not load, so this belongs to the technical team, but the word "billing" pulls toward billing.
We built two sets of 200 pairs for this, written by two different language models and checked by us. In each pair, one case has such a lure and the other uses the same word where it really points to the right option. The table shows how often both cases of a pair are right, and in brackets how often the model picked the lure. With descriptions, every option also came with a short description, as in the example above.
| model | set A | set A, with descriptions | set B | set B, with descriptions |
|---|---|---|---|---|
| Onebox 1 (9B) | 85.0 % (9.5 %) | 90.0 % (7.5 %) | 60.5 % (25.0 %) | 75.5 % (19.5 %) |
| Clef-Flash (9B) | 90.5 % (6.0 %) | 94.5 % (5.5 %) | 69.5 % (25.5 %) | 83.0 % (14.5 %) |
Clef-Flash is clearly better here. The hardest case for both is a broken page, form or app that is named after a topic, like the billing page above. On the billing page itself, every model we tried picked billing (Onebox 99.7 %, Laya 99.9 %, Julia-1 98.9 %), except Clef-Flash, which picked technical with 70 %. This is the main thing we work on for the next version.
On our 104 older lure cases, Onebox falls for 7 (openjev 4B: 18, openjev 0.8B: 57).
Other tasks
150 items per task, fixed sample, none of them seen in training.
| task | options | Onebox 1 (9B) |
|---|---|---|
| AG News | 4 | 90.0 % |
| emotion | 6 | 54.7 % |
| XNLI (English) | 3 | 83.3 % |
| Banking77 | 77 | 55.3 % |
| MASSIVE | 59 | 70.0 % |
Speed
With onebox.py in bf16 on a MacBook Pro with M5 Max (Apple MPS), a question takes about 0.9 seconds. On the Mac, PyTorch only has slow reference kernels for the linear attention layers of Qwen3.5. Our own MLX server for Apple Silicon answers the same questions in 0.33 seconds on average (0.27 median) and in 0.31 seconds with an 8-bit base model (9.7 GB, 98.3 % same answers). That server is not part of this release yet. We have not measured onebox.py on an NVIDIA GPU.
Questions with many options stay cheap because the shared part of the input is read once. With 77 options (Banking77) a question takes 0.64 seconds on the MLX server.
Training data
- Public sets with their own labels: SNLI, ANLI, MultiNLI, BoolQ.
- Decision cases we wrote or generated: workflow decisions, implicature and sarcasm, ordinal scales, and training cases built against lures. None of these overlap with the lure test sets.
- The training split of typed-decisions (5370 decisions), never the test split.
Training ran in one stage on one rented H100 (about 2.3 hours, 7200 steps of 4 decisions). The numbers on this card were measured on a Mac with MLX. We checked that the released PyTorch code gives the same decision as those measurements on 398 of 400 test decisions (the rest is rounding in bf16).
Limits
- Words that name an option still pull Onebox toward it, most of all when a page, form or app named after a topic is broken. See the table above.
- On workflows it has not seen, expect less than 79.3 %. The same recipe without the typed-decisions training split reaches 59.4 %.
- Mostly English. We have not measured other languages.
- It was trained on inputs of up to 768 tokens (about 550 English words).
onebox.pyreads up to 32,000 tokens and cuts longer states from the beginning. How well it decides on longer inputs is not measured yet. - It reads text only, no images.
- It is a research model. Do not use it for decisions about people without a human checking the result.
License
The weights and onebox.py are released under the Apache License 2.0. The base model Qwen3.5-9B is licensed under Apache 2.0 by the Qwen team.
Citation
@misc{newcombs2026onebox,
title = {Onebox 1: typed decisions in one pass},
author = {{newcombs}},
year = {2026},
url = {https://huggingface.co/newcombs/onebox-1-9b}
}