diffcider-typed-decisions
A typed-decision model made with sysone: the encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 (lora) under a 0-layer yesno head, trained on LocalLLaMA/typed-decisions (all) at d0e2f0c4.
The adapter makes the masked diffusion language model Qwen3-0.6B-diffusion-mdlm-v0.1 a typed-decision model: it answers choices, scores and yes/no statements in one forward pass, reading the model's own Yes against its No at a mask after each option, in jev-dllm's layout. On typed-decisions' 2,000 test decisions it scores 0.768, where the model scores 0.409 untrained and jev-dllm reports 0.5225 for a full fine-tune of it. With the adapter switched off the model is the published one exactly, and it writes (dec.generate(prompt)). sysone's mdlm module loads the checkpoint's a2d-qwen3 model type with its own classes, so no remote code runs.
The repo holds the head and a LoRA adapter, not the encoder's weights: Decider.load downloads dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 at commit c8d24a3f4a from the Hub and puts the adapter on it.
Use it
Install sysone from GitHub, in Python 3.11 or newer:
pip install git+https://github.com/sgaseretto/sysonelib
A state is whatever the decision is about, as JSON-like data. Each question has a type: a choice among named options, a score on an ordered scale, or a noul, a statement that is true or false. The options are named when asking, so they can be new ones.
from sysone.inference import Decider
dec = Decider.load("sgaseretto/diffcider-typed-decisions")
state = {
"account": {"tier": "premium", "tenure_months": 26, "prior_tickets_90d": 2},
"thread": [
{
"role": "customer",
"text": "I was charged twice for my March invoice, and nobody has answered my last two emails.",
},
{"role": "agent", "text": "I'm sorry about that. Could you share the invoice number so I can look into it?"},
{"role": "customer", "text": "It's INV-3381. Refund the second charge today or I'm cancelling the account."},
],
}
questions = {
"category": {
"type": "choice",
"instructions": "What is this customer conversation primarily about?",
"criteria": {
"billing": "A charge, invoice, subscription or payment problem.",
"refund": "The customer is explicitly asking for money back.",
"technical": "The product or service is not working as expected.",
},
},
"urgency": {
"type": "score",
"instructions": "How time-sensitive is this conversation?",
"criteria": [
"No time pressure; can wait indefinitely.",
"Routine; handle within the normal queue.",
"Elevated; should be handled within the same week.",
"Critical; requires action within the same day.",
],
},
"needs_human": {
"type": "noul",
"instructions": "This conversation requires a human agent rather than automated handling.",
},
}
answers = dec.predict(state, questions)
answers holds an answer per question, in the Jev schema: a choice names the likeliest option and gives every option's probability, a score gives its expected level on the scale (0 for the first) and every level's probability, and a noul gives the probability that its statement is true; each says how confident it is. This model's answers to the example:
{
"category": {
"type": "choice",
"choice": "refund",
"probabilities": {"billing": 0.429, "refund": 0.5677, "technical": 0.0033},
"confidence": 0.3597,
"answer_confidence": 0.5677,
},
"urgency": {
"type": "score",
"score": 2.6442,
"legend": {
"0": "No time pressure; can wait indefinitely.",
"1": "Routine; handle within the normal queue.",
"2": "Elevated; should be handled within the same week.",
"3": "Critical; requires action within the same day.",
},
"probabilities": {"0": 0.0069, "1": 0.0483, "2": 0.2387, "3": 0.7062},
"confidence": 0.4459,
"answer_confidence": 0.7062,
},
"needs_human": {"type": "noul", "noul": 0.7083, "confidence": 0.7083, "answer_confidence": 0.7083},
}
Results
On LocalLLaMA/typed-decisions's test split (2,000 decisions), calibrated, on a Kaggle T4:
| metric | value |
|---|---|
| accuracy | 0.7675 |
| accuracy_choice | 0.7517 |
| accuracy_score | 0.7238 |
| accuracy_noul | 0.8417 |
| ece | 0.1496 |
| brier | 0.0536 |
On the validation split (400 decisions), during training:
| metric | value |
|---|---|
| loss | 0.9801 |
| accuracy | 0.6925 |
| soft_accuracy | 0.4572 |
| brier | 0.0621 |
| kl | 0.1006 |
| tv | 0.167 |
| ece | 0.1495 |
| score_mae | 0.2234 |
| within_one | 1 |
| accuracy_choice | 0.7 |
| accuracy_score | 0.6062 |
| accuracy_noul | 0.85 |
Calibration temperatures: choice 0.945, choice:3-5 0.945, score 1.02, score:3-5 1.02, noul 1.02, noul:2 1.02.
Training
| Data | LocalLLaMA/typed-decisions (all) at d0e2f0c4; 1080 train, 80 valid, 120 calib cases |
| Encoder | dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1, lora (LoRA r=16, α=32, on down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj) |
| Head | 0 layers, yesno scorer, readout anchor |
| Loss | soft_ce_rps=0.25, preset t4, seed 0 |
| Fit 1 | 1 epoch, cosine, lr 0.001 (encoder 0.0002), batch 8×4, fp16 on cuda: 300 steps in 87.4 min, peak memory 7.69 GB, final loss 0.878 |
| Machine | Intel(R) Xeon(R) CPU @ 2.00GHz, 4 cores, 31.3 GB; accelerator cuda (Tesla T4, Tesla T4); a kaggle run (job mdlm-lora-check) |
Reproduce
- sysone:
sysonelibat commit123ad8bac6(main), with uncommitted changes in 22 files (diff SHA-25677eb5a356ea0) - Ran:
run_b.py - Python 3.13.15, sysone 0.3.0, torch 2.11.0+cu128, transformers 5.16.1, peft 0.20.0, accelerate 1.14.0, datasets 4.8.5, huggingface_hub 1.29.0, tokenizers 0.23.1, safetensors 0.8.0, numpy 2.1.3, fastcore 2.2.32, plum-dispatch 2.10.1;
environment.txtlists every package
Model tree for sgaseretto/diffcider-typed-decisions
Base model
Qwen/Qwen3-0.6B-BaseDataset used to train sgaseretto/diffcider-typed-decisions
Evaluation results
- loss on typed-decisionsself-reported0.980
- accuracy on typed-decisionsself-reported0.693
- soft_accuracy on typed-decisionsself-reported0.457
- brier on typed-decisionsself-reported0.062
- kl on typed-decisionsself-reported0.101
- tv on typed-decisionsself-reported0.167
- ece on typed-decisionsself-reported0.149
- score_mae on typed-decisionsself-reported0.223