Instructions to use sthanika-ai/sieve-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sthanika-ai/sieve-2b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Sieve-2B
Sieve is a decision model: one state, one typed question, a closed set of options in -> a calibrated probability distribution over the options out. No text generation, no sampling.
It is a LoRA adapter plus a small scalar scorer on Qwen/Qwen3.5-2B-Base, truncated to the base
model's 16th of 24 blocks:
| parameters | |
|---|---|
| run at inference (embeddings + blocks 1-16 of the base, + adapter + head) | 1.43 B |
| trained by us (LoRA 3,637,248 + scorer head 1,053,697) | 4.69 M |
full base model (Qwen/Qwen3.5-2B-Base text model), for reference |
1.88 B |
Blocks 17-24 and the base model's text-generation head never run, so they are not loaded onto the GPU.
Not affiliated with TypeSafe AI, Jev, or kev. Sieve is an independent project; comparisons
below are against jaredpalmer/kev-4b (a different, larger open decision model on the same base
architecture) as a public reference point, not an endorsement or partnership.
What makes it different: order-invariant by construction
Most decision models that read several options in one pass (kev included) can change their answer when the same options are given in a different order. kev's own project documentation reports an 11.3% argmax-flip rate under option reshuffling. Sieve cannot flip: the state and question are encoded once, that encoded state is copied once per option, and each option is scored as a clean continuation of its own copy, never seeing its siblings. This is a property of the architecture, not something learned: it holds by construction, and was verified in development at a numerical precision of ~1e-6 and as exactly 0 argmax flips across 600 real banking77 examples under option reshuffling.
from modeling_sieve import Sieve # needs: transformers>=5.17, torch, peft, safetensors, huggingface_hub
s = Sieve.from_pretrained("sthanika-ai/sieve-2b") # downloads Qwen3.5-2B-Base once (~4.3 GB), applies this LoRA + head
s.choice(
state="Customer: my invoice was charged twice and nobody answers the phone!",
question="Which team should handle this?",
options=["Billing", "Technical support", "Sales"],
)
# -> {"decision": "Billing", "index": 0, "confidence": 0.81,
# "probs": {"Billing": 0.81, "Technical support": 0.16, "Sales": 0.03}}
Reordering options gives back the identical distribution (only the dict order changes), and
the probs field is a genuine probability distribution over the options actually given.
Training
- Base:
Qwen/Qwen3.5-2B-Base, a hybrid of Gated DeltaNet (linear-attention) and full-attention blocks. Only its first 16 of 24 blocks ever run (read_layer=16). - Adapter: LoRA, r=8, α=16, on the attention and MLP projections of blocks 1-16 (3,637,248 parameters), plus a scalar scorer head (LayerNorm -> 512 -> 1, 1,053,697 parameters).
- Recipe: from scratch, 2 epochs, lr 5e-5, training prompts up to 2,048 tokens. An earlier run at lr 2e-4 (4x higher) scored 13.8 points lower on MMLU with everything else held constant -- evidence that a lower rate preserves more of the base model's own knowledge during LoRA fine-tuning. See "Caveats": this finding is from a single seed pair.
- Data, per epoch (26,241 rows):
- ~18,200 rows of kev's decision-v7 public sources (banking77, boolq, ag_news, multi_nli, sst5, yelp_review_full, trec, dbpedia_14, amazon_reviews_multi_en, imdb; listed above), with kev's own policy-rule and contrast examples;
- 7,999 synthetic option-comparison questions generated by us (pick the largest / second-largest / median / closest-to-target / sum of a set of numbers or expressions). No Jev outputs were used for training or labels.
- The prompt lists every option, in an order fixed by a hash of each option's own text -- the order depends only on which options are present, never on their order or value, which is what keeps the invariance above exact while still letting the model see the whole option set.
Results (this checkpoint, single seed)
| benchmark | Sieve-2B | kev-4b* |
|---|---|---|
| MASSIVE (intent, K=59) | 71.5 | 70.2 |
| MMLU | 60.8 | 70.5 |
| SciQ | 97.5 | 98.7 |
| transfer-v4 (clean, out-of-domain) | 73.2 | 87.8 |
| held-out pool (unseen schemas, K up to 150) | 77.7 | 76.5** |
| banking77 (in-domain for both) | 76.5 | 78.3 |
| answer changes under option reordering | 0.0% | ~11.3%*** |
*kev-4b is 4B parameters, roughly double this model, and trains on a more mature multi-round recipe; this table compares a 2B model against a 4B one, not a like-for-like size. **kev-4b rejects high-option-count items past its own token budget and they count as wrong here; Sieve has no such limit. ***kev's own reported figure (see above); not independently remeasured here.
Decision Index 0.2.1 (self-run with the official kit; not yet reviewed by the board)
Scored with the Decision Index reproduction kit on
the full frozen suite (rebuilt byte-identical to the maintainers' files: selected rows sha256
b2b56d6f..., added rows 7429f3c9...): 150,317 scored requests, complete, coverage 1.00 in
every area.
| Decision Index | raw | breadth | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Human Taste | |
|---|---|---|---|---|---|---|---|---|
| Sieve-2B | 21.74 | 39.91 | 20.55 | 10.6 | 19.7 | 28.5 | 35.3 | 17.8 |
- Files:
decision_index/in this repo --scores.json,index.json,benchmark-summary.json,environment.json,status.json, andresults.jsonl.gz(the kit's--compactform: payloads dropped so no suite text is republished; every row keeps itspayload_sha256, and the file re-scores to the identical index). - Engine:
sieve_engine.pyin this repo (--engine sieve_engine:SieveEngine), bf16, context limit 32,768 tokens, optionsmem_fraction=0.16,max_batch_tokens=49152(these options reproduce the recorded probabilities bit-for-bit; defaults give the same decisions). 25 requests refused (two options with identical text); 0 errors. - Execution: 24 shards of exactly the rows
pipelineruns, 7 concurrent processes on 2x A100 80GB, merged by run_id (details inenvironment.json). Single-process latency: median 85 ms, p95 223 ms (400 uniformly sampled suite requests, A100). - Overlap with the suite: we checked every training row and the temperature-fitting pool against every suite row (inputs only: a shared 13-word run, or a whole 5-12-word field contained in the other side). 33 of the 150,317 scored rows (0.02%) share an input with the training data: BANKING77 10 (short customer queries that occur word-for-word in BANKING77's own train and test splits), HoVer 9 and RAGTruth 6 (shared source sentences with unrelated labels), ANLI 2, SATA-Bench 2, iSarcasmEval 2, BRIGHT 1, MMLU 1 (not in the index). No benchmark's test set was used for training. The temperature pool (used only to fit the single temperature, which never changes a decision) shares inputs with 85 suite rows, mostly ContractNLI clauses, CLINC150 utterances and MuSR. Every run_id is listed in
decision_index/overlap_run_ids.json.
Calibration
Calibration is built in, with no labels needed from you -- a single shipped temperature, fit
once by us on data this checkpoint never trained or was tested on (a held-out pool of unseen
schemas: clinc, newsgroups, yahoo, LEDGAR), and applied automatically by choice().
config.json's "temperature": 1.7144 is that value; loading the model already gets you this.
Measured effect on held-out benchmarks (Expected Calibration Error, lower is better; accuracy is provably unchanged, since a temperature only rescales confidence, never the decision):
| benchmark | raw (T=1.0) | built in (T=1.71) |
|---|---|---|
| MMLU | 0.135 | 0.040 |
| transfer-v4 (clean) | 0.114 | 0.074 |
| transfer-v4 (all) | 0.122 | 0.049 |
| kev's own calibration split | 0.050 | 0.032 |
| SciQ | 0.013 | 0.018 |
| banking77 | 0.079 | 0.082 |
| MASSIVE | 0.088 | 0.094 |
4 of 7 benchmarks improve, with the largest gains on MMLU and transfer-v4. SciQ, banking77 and MASSIVE get slightly worse -- one shared number can't correct every benchmark's own direction of miscalibration at once. If your use case is MASSIVE- or banking77-shaped specifically, the per-schema refit below will do better than either column here.
Optional, sharper upgrade if you have your own labels: call fit_schema(name, examples) once
with ~100 labelled (state, question, options, gold_index) examples for your specific kind of
question, then pass schema=name to choice(...). This is a 2-parameter logistic refit of the
confidence (never the decision itself):
| benchmark | raw (T=1.0) | built in (T=1.71) | with 100-label fit_schema |
|---|---|---|---|
| MASSIVE | 0.089 | 0.094 | 0.059 |
| banking77 | 0.078 | 0.082 | 0.050 |
| MMLU | 0.134 | 0.040 | 0.063 |
| transfer-v4 (clean) | 0.116 | 0.074 | 0.072 |
| kev's own calibration split | 0.049 | 0.032 | 0.042 |
| pool (unseen schemas) | 0.087 | -- | 0.041 |
fit_schema is never required -- the model is fully usable, and already reasonably calibrated,
without it.
Context length
config.json sets max_len: 32768: inputs up to 32,768 tokens are scored; longer ones are trimmed
from the front of the state. Training used prompts up to 2,048 tokens. Long-context accuracy was
measured on Sieve-4B, not this model: on MMLU with 4K-30K tokens of unrelated text prepended,
accuracy dropped by a flat ~5-8 points with no further decline from 4K to 30K.
Caveats
- Single seed. Earlier multi-seed runs of the same recipe family showed a spread of roughly +/-0.2-1.5 points on real benchmarks; this exact checkpoint has not been replicated on a second seed.
- The lr 5e-5 vs 2e-4 finding is from one paired comparison, not a multi-seed sweep.
- MMLU and transfer-v4 lag a larger reference model (kev-4b, 4B parameters) on general knowledge and kev's own out-of-domain suite; this is the known, open gap for this model size.
- Requires
transformers>=5.17for theqwen3_5architecture andpeft>=0.21.
How it was built
The prefix-state fork (see modeling_sieve.py) replicates DynamicCache -- both the attention KV
cache and the Gated DeltaNet layers' recurrent state -- once per option, then runs every option as
one right-padded batch of continuations. Sieve.from_pretrained defaults to fp32 with TF32 matmul.
The Decision Index engine runs bf16 for speed; on a Sieve-4B sample, bf16 changed 1 of 523
decisions (0.19%) relative to fp32.
License
Apache-2.0 for the adapter, head and code in this repository. The Qwen3.5 base model is
Apache-2.0. The public training datasets each keep their own license (see the datasets: field
above and each dataset's own card); several of kev's decision-v7 sources (ag_news, yelp_review_full,
imdb, sst5, trec, amazon_reviews_multi) do not state a licence on their Hub cards -- check their
original terms before commercial use.
- Downloads last month
- -
Model tree for sthanika-ai/sieve-2b
Base model
Qwen/Qwen3.5-2B-BaseDatasets used to train sthanika-ai/sieve-2b
fancyzhx/ag_news
google/boolq
Spaces using sthanika-ai/sieve-2b 2
Collection including sthanika-ai/sieve-2b
Evaluation results
- accuracy on MASSIVE, held out (1,200 questions)self-reported0.715
- expected_calibration_error on MASSIVE, held out (1,200 questions)self-reported0.094
- accuracy on MMLU, held out (1,500 questions)self-reported0.608
- accuracy on transfer-v4 clean (328 questions, never trained on)self-reported0.732
- accuracy on held-out pool (clinc, newsgroups, yahoo, LEDGAR)self-reported0.777