Instructions to use sthanika-ai/Sieve-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sthanika-ai/Sieve-9B with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "sthanika-ai/Sieve-9B") - Notebooks
- Google Colab
- Kaggle
Sieve-9B
Sieve-9B is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.
It is a LoRA adapter (rank 32) and a 2.1M-parameter pointer head on Qwen/Qwen3.5-9B. The code and training data are at github.com/sthanika-ai/sieve.
How it works
Sieve follows the decision-model design of Kev:
- Backbone. Only the text part of Qwen3.5-9B is kept: 32 layers and the embeddings, 7.94B parameters.
- Removed: the language-model head, the multi-token-prediction layer and the vision encoder, 1.72B parameters together.
- Readout. A pointer head scores the end of each option against the decision point.
- Isolation. Each question runs as its own row from the state's cache, so questions never see each other.
What Sieve-9B changes:
- Backbone. The post-trained Qwen3.5-9B instead of the base model.
- Averaged weights. The adapter and head are the exact average of two rank-16 runs from one initialisation.
- New data. 1,920 records from 240 rule structures not in the decision-v7 split, plus states padded with unrelated text.
- Serving:
- CUDA graphs: 2 options take 18.6 ms instead of 70.2 ms.
- Long questions run as one row: 255 options take 330.5 ms instead of 371.0 ms.
| type | options | returns |
|---|---|---|
choice |
the caller's keys, optionally described (up to 255) | a probability per key |
noul |
yes / no | P(yes) |
score |
ordered levels | a probability per level |
Use
pip install "sieve-decisions[cuda,serve] @ git+https://github.com/sthanika-ai/sieve"
from sieve import load_sieve, decide
m = load_sieve("sthanika-ai/Sieve-9B")
decide(m, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
{"department": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
"churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
- Loading. Load it with
sieve; the adapter alone is not a text-generation model. - Server.
sieve-serve --model sthanika-ai/Sieve-9B --graphs. - GPU memory. About 20 GB; 17.6 GiB was measured with a 4,000-token state.
- Temperature.
head.ptstores T = 1.631, fitted by NLL on the decision-v7 calibration split (1,148 questions).- That split is used only to fit T and to choose the checkpoint; no reported number comes from it.
- T never changes an answer, and
temperature=1.0gives raw probabilities.
Results
All numbers are Sieve-9B's own. They are recomputed from its saved per-question predictions and stored in
result.json.
sieve-eval in the GitHub repository computes the same metrics for any
labelled JSONL file. With --no-merge it reproduces the public-test rows below exactly.
Decision Index 0.2.1
Sieve-9B scores 41.71 (raw index 56.47, breadth 40.06) on our own run of the full frozen
Decision Index suite. The run used the official
kit at 87d4650 and its scorer.
- Coverage. 150,317 requests over 44 benchmarks, all answered, with no truncation or errors. The index averages 38 of the benchmarks.
- Results. sthanika-ai/Sieve-9B-decision-index-results.
- Not on the leaderboard yet. The score is self-reported and expected, not confirmed, until the maintainers review it. On the 2026-09-28 board it would place about 18th of 72.
| Decision Index | Knowledge & Reasoning | Language Understanding | Retrieval & Classification | Tools & Automation | Arts & Human Taste |
|---|---|---|---|---|---|
| 41.71 | 27.4 | 46.1 | 46.1 | 61.3 | 22.8 |
The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. The per-benchmark scores below are in each benchmark's own metric.
All 44 benchmarks
| benchmark | metric | score | requests |
|---|---|---|---|
| ACOS | per-review F1 | 0.241 | 1,565 |
| ANLI | macro-F1 | 0.562 | 3,200 |
| API-Bank | accuracy | 0.726 | 508 |
| Amazon ESCI | macro-F1 | 0.493 | 5,000 |
| BANKING77 | macro-F1 | 0.836 | 3,080 |
| BBH fixed-option tasks | accuracy | 0.672 | 5,507 |
| BFCL | case exact accuracy | 0.954 | 1,694 |
| BPoMP | accuracy | 0.665 | 5,000 |
| BRIGHT | nDCG@10 | 0.420 | 220 |
| CLINC150+OOS | macro-F1 | 0.760 | 5,500 |
| CLadder | accuracy | 0.633 | 5,000 |
| CRUXEval | accuracy | 0.537 | 570 |
| ChessBench | accuracy | 0.114 | 5,000 |
| ContractNLI | macro-F1 | 0.615 | 123 |
| FinEntity | macro-F1 | 0.888 | 979 |
| ForecastBench | Brier (lower is better) | 0.181 | 10,139 |
| GPQA Diamond | accuracy | 0.424 | 196 |
| GSM8K | accuracy | 0.460 | 2,638 |
| HLE | accuracy | 0.106 | 501 |
| Habermas Machine | accuracy | 0.404 | 1,676 |
| HellaSwag | accuracy | 0.830 | 10,042 |
| HoVer claim verification | accuracy | 0.635 | 4,000 |
| Home appliance simulator | case exact accuracy | 0.307 | 88 |
| Humicroedit | accuracy | 0.551 | 2,628 |
| MMLU-Pro | accuracy | 0.508 | 12,032 |
| MuSR | accuracy | 0.567 | 752 |
| NLI4CT | macro-F1 | 0.779 | 5,500 |
| New Yorker caption matching | accuracy | 0.581 | 528 |
| POP909-CL | accuracy | 0.129 | 2,000 |
| PhishNChips phishing decisions | accuracy | 0.541 | 2,000 |
| RAGTruth response-level hallucination | F1 on hallucinated class | 0.629 | 2,700 |
| SATA-Bench | case exact accuracy | 0.296 | 1,650 |
| ToolRet | nDCG@10 | 0.651 | 685 |
| VAST | macro-F1 | 0.522 | 3,006 |
| When2Call MCQ | accuracy | 0.561 | 3,652 |
| WinoGrande | accuracy | 0.751 | 1,267 |
| cfcolor | accuracy | 0.580 | 5,000 |
| iSarcasmEval | Sarcasm F1 · track A, English | 0.506 | 4,600 |
| ARC-Challenge (not in the index) | accuracy | 0.938 | 1,172 |
| ARC-Easy (not in the index) | accuracy | 0.978 | 2,376 |
| MMLU (not in the index) | accuracy | 0.749 | 14,033 |
| RouterBench (not in the index) | selected quality (quality objective) | 0.793 | 10,000 |
| SGD/SGD-X (not in the index) | macro-F1 | 0.627 | 2,500 |
| SimpleBench (not in the index) | accuracy | 0.200 | 10 |
Decision suites, held-out benchmark and public test sets
| benchmark | questions | accuracy [95% CI] | NLL | Brier | ECE | wrong at p ≥ 0.9 |
|---|---|---|---|---|---|---|
| decision-v7 development (in-domain) | 1,264 | 0.872 [0.851, 0.892] | 0.330 | 0.176 | 0.028 | 1.7% |
| transfer-v4 development (out-of-domain) | 656 | 0.829 [0.797, 0.860] | 0.447 | 0.242 | 0.045 | 3.7% |
| transfer-v9 development, strict (out-of-domain) | 936 | 0.753 [0.724, 0.781] | 0.689 | 0.337 | 0.050 | 3.8% |
| held-out decision benchmark ³ | 6,224 | 0.857 [0.848, 0.865] | 0.451 | 0.207 | 0.055 | 0.4% |
| MMLU, 57 subjects, 4-way ² | 14,042 | 0.758 [0.751, 0.765] | 0.642 | 0.333 | 0.008 | 1.7% |
| SciQ | 1,000 | 0.984 [0.976, 0.991] | 0.063 | 0.026 | 0.024 | 0.1% |
| Banking77, 77 intents ¹ | 3,079 | 0.857 [0.844, 0.869] | 0.563 | 0.224 | 0.043 | 1.6% |
| MASSIVE intent, 60 labels | 2,974 | 0.765 [0.749, 0.780] | 0.886 | 0.345 | 0.093 | 0.5% |
| MASSIVE scenario, 18 labels | 2,974 | 0.711 [0.695, 0.728] | 0.889 | 0.401 | 0.059 | 0.2% |
| typed-decisions test | 2,000 | 0.704 [0.681, 0.726] | 0.704 | 0.404 | 0.040 | 0.5% |
- Calibration. The probability metrics use the stored temperature (T = 1.631), as served. Raw (T = 1) values are
in
result.json. - Wrong at p ≥ 0.9 is the share of all questions answered wrongly with a probability of at least 0.9.
- Intervals. They come from a bootstrap with 4,000 resamples over groups of related records.
- Development partitions only. The decision suites are development partitions; no test partition of them has been scored.
- Unmerged adapter. These rows were scored with the adapter unmerged. The Decision Index run and the latency use the merged model, as served.
¹ In-distribution: the training data contains 1,000 Banking77 training questions.
² Our conversion. On the Decision Index's own MMLU, 14,033 questions outside the index, Sieve-9B scores 0.749.
³ The held-out benchmark:
- Size. 6,233 questions over 10 synthetic decision domains, with no family overlap with training. 6,224 were scored; 9 were dropped because their options exceed the evaluation's 12,288-token question limit.
- Truncation. States over 384 tokens, the training limit, were truncated to 384.
- Option counts. 77 of the scored questions have 256 or 512 options, more than the served API accepts (255).
| held-out benchmark by | questions | accuracy |
|---|---|---|
| state ≤ 384 tokens | 5,186 | 0.889 |
| state > 384 tokens (truncated) | 1,038 | 0.697 |
| 2 options | 2,452 | 0.976 |
| 3–10 options | 2,969 | 0.823 |
| 11–50 options | 522 | 0.728 |
| 51–255 options | 204 | 0.431 |
| 256–512 options | 77 | 0.377 |
Tool selection (BFCL v3, decision form)
The model picks the function to call from the offered ones, or "none of these". Argument values are not generated or scored, so these are not BFCL leaderboard numbers.
| BFCL v3, decision form | questions | accuracy |
|---|---|---|
| function selection (simple, multiple, live simple, live multiple) | 1,909 | 0.963 |
| irrelevance detection (irrelevance, live irrelevance) | 1,118 | 0.720 |
Latency
Server-reported p50 on one A100 80GB (bf16, merged, CUDA graphs), for one choice question on a cached state of about 75 tokens, excluding tokenisation and HTTP:
| options | 2 | 5 | 10 | 25 | 50 | 100 | 255 |
|---|---|---|---|---|---|---|---|
| ms | 18.6 | 19.3 | 22.8 | 42.8 | 74.5 | 128.4 | 330.5 |
Training
- Recipe:
- LoRA r=16, α=32 on the attention, MLP and DeltaNet projections, with the pointer head trained from scratch;
- lr 5e-5, OneCycle, 2 epochs, 8 records per step, bf16;
- permuted choice options and "none of the above" augmentation, with a cross-entropy loss.
- Averaging. Two adapters were trained from one initialisation on different data mixes and averaged 50/50 ("model soups", Wortsman et al., 2022).
- Selection. The checkpoint was chosen by calibration-split accuracy, then NLL. That rule was fixed after both
runs were evaluated and before the average was scored.
- The benchmarks were never used to choose. The second data mix did target weaknesses seen on the development partitions.
- No outputs of other models were used.
- Data. The four training files are in
data/sieve-9b/, unchanged.- The recipe and hashes are in
training_config.json, and the logs intrain.log. - Sources and licences are in the repository's NOTICE.
- The recipe and hashes are in
- Data hygiene. The training files were checked against seven benchmark partitions (listed in
provenance.json).- No overlap: by exact question, exact state, and an order-insensitive hash of each state.
- Near-duplicates: the unknowable date-policy file has 4 near-duplicate states, and its families are excluded from every transfer-v9 number.
- Rule structures: 30 of the 240 are logically equivalent to a decision-v7 structure after pushing negations inward. None equals or contains a held-out structure.
Files
| file | contents |
|---|---|
adapter_model.safetensors, adapter_config.json |
the LoRA adapter (rank 32, PEFT format) |
head.pt |
the pointer head and the calibrated temperature |
tokenizer.json, tokenizer_config.json |
Qwen3.5-9B's tokenizer, unchanged |
training_config.json, training_metrics.json, train.log |
the recipe, the two runs and how they were averaged |
result.json |
the numbers above, from saved per-question predictions, with the same metrics as sieve-eval |
provenance.json |
SHA-256 of every file and of the training data, and library versions |
Credits and license
- Design. The typed-decision format, the pointer-head design and the stored temperature follow Kev (Apache-2.0).
- Training data. It comes from Kev's decision-v7 suite, its date-policy files and its rule generator. It also contains text from ten public datasets under their own licences, listed in the GitHub NOTICE.
- Backbone. Qwen3.5-9B, Apache-2.0.
- License. The weights are Apache-2.0 (see
LICENSEandNOTICE).
- Downloads last month
- 59
Model tree for sthanika-ai/Sieve-9B
Datasets used to train sthanika-ai/Sieve-9B
fancyzhx/ag_news
google/boolq
Spaces using sthanika-ai/Sieve-9B 2
Collection including sthanika-ai/Sieve-9B
Evaluation results
- Decision Index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.2.1 (150,317 requests over 44 benchmarks; the index averages 38)self-reported41.710
- accuracy on decision-v7 development (1,264 clean questions)self-reported0.872
- ECE, at the shipped temperature on decision-v7 development (1,264 clean questions)self-reported0.028
- accuracy on transfer-v4 development (656 clean questions)self-reported0.829
- Brier, at the shipped temperature on transfer-v4 development (656 clean questions)self-reported0.242
- accuracy on transfer-v9 development, strict (936 questions)self-reported0.753