Instructions to use FINAL-Bench/Darwin-27B-ZTC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/Darwin-27B-ZTC with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="FINAL-Bench/Darwin-27B-ZTC")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("FINAL-Bench/Darwin-27B-ZTC") model = AutoModel.from_pretrained("FINAL-Bench/Darwin-27B-ZTC", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Darwin-27B-ZTC
A zero-token decision engine from the Darwin family. Darwin-27B-ZTC reads a piece of state and a set of typed questions, and returns a full probability distribution over the options of every question. It does this in one forward pass per question. It generates no tokens.
It takes the POST /v1/systemone request shape: noul (yes or no), choice (one of N labels) and score (an ordered rubric). Every option can carry a written description, and the descriptions are part of the input.
Results
Typed Decisions (general, zero-shot)
LocalLLaMA/typed-decisions, test split: 400 cases, 2,000 decisions. One request per case, with the state and all five questions together, the same request shape as the benchmark README.
| Metric | Darwin-27B-ZTC |
|---|---|
| Accuracy ↑ | 0.743 |
| KL from gold ↓ | 0.204 |
| Brier ↓ | 0.097 |
| Question type | Decisions | Accuracy |
|---|---|---|
noul (yes or no) |
600 | 0.845 |
choice |
600 | 0.732 |
score (rubric) |
800 | 0.675 |
Zero-shot. The model never saw the Typed Decisions train split, its four workflows or its question schemas. All 2,000 decisions were answered, with zero errors.
Scoring follows the benchmark README: accuracy is agreement with the gold label, KL is KL(gold || prediction) averaged over decisions, and Brier is the squared error summed over the options of a decision, averaged over decisions. Our scorer reproduces the README's Uniform reference (KL 0.444, Brier 0.238). We ran the test set twice with the same model and settings. The two runs gave accuracy 0.741 and 0.743. Every number on this card comes from the second run, the one with the saved predictions (noul 507/600, choice 439/600, score 540/800, 1,486/2,000 in total). The two runs differ on 14 of 2,000 decisions (9 better, 5 worse; mean +0.002, 95% case-bootstrap interval -0.002 to +0.006), which is within run-to-run noise: the server batches questions from concurrent requests, so near-tie decisions can flip between runs. An earlier version of this card mixed the per-type numbers of the first run with the headline of the second; thanks to the community member who caught it.
How it works
- Backbone: FINAL-Bench/Darwin-27B-RSI, full-weight trained as a decision engine.
- Readout: the options are listed in the prompt with short codes. A readout head maps the final hidden state of one forward pass to one score per option. A softmax over the scores gives the distribution, and a fitted scalar temperature calibrates it.
- Training data: public training and development splits of public decision benchmarks only. No test split of any benchmark was used for training.
Files
| File | What it is |
|---|---|
model-*.safetensors, model.safetensors.index.json, config.json |
backbone weights (BF16) |
readout.safetensors |
decision readout head |
decision_config.json |
option codes, calibration temperature and provenance |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
tokenizer and prompt template |
Usage
This repo ships the inference code in autojev/: the model.py and types.py of the autojev decision-model code (MIT, copyright notice in autojev/LICENSE), with support for a text-only backbone added. It needs torch, transformers (Qwen3.5 support), safetensors and pillow, and a GPU with space for about 54 GB of BF16 weights.
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("FINAL-Bench/Darwin-27B-ZTC")
sys.path.insert(0, path)
from autojev.model import DecisionModel
model = DecisionModel(checkpoint=path)
row = {
"state": {"ticket": "Customer was charged twice for the same order."},
"question": {
"type": "choice",
"instructions": "What should support do?",
"criteria": {"refund": "Refund the duplicate charge.", "escalate": "Send to billing.", "close": "No action."},
},
}
probabilities = model.predict([row])[0] # one probability per option, in criteria order
Citation
@misc{darwin27bztc2026,
title = {Darwin-27B-ZTC: a zero-token decision engine},
author = {VIDRAFT and FINAL-Bench},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC}
}
- Downloads last month
- 2
Collections including FINAL-Bench/Darwin-27B-ZTC
Article mentioning FINAL-Bench/Darwin-27B-ZTC
Evaluation results
- LocalLLaMA/typed-decisions leaderboard
- Accuracy View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.0.74 * - Kl From Gold View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.0.2 * - Brier View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.0.1 *