buridan-2b

A decision model: it picks one of the options it is given and says how sure it is, in one forward pass, without generating text. A replication of Strands Decider on Qwen/Qwen3.5-2B-Base. buridan.py in this repo trains, evaluates and loads it.

Named after Buridan's ass, the donkey that starved between two equal piles of hay because it could not choose. This one always chooses. When the piles look equal, its confidence says so.

A donkey stands between a hay bale and a water trough

🚀 Usage

uv run https://huggingface.co/jgalego/buridan-2b/resolve/main/buridan.py ask \
  --state "Help! My payouts have been failing for 3 days!" \
  --choice "Which team should handle this?=billing,sales,retail" \
  --noul "Does this convey urgency?" \
  --score "How frustrated is the writer?=calm,frustrated,depressed"

Three question types share one head:

Type Options Returns
noul false, true P(true)
choice 2 or more, set per request the best option, every option's probability, a confidence
score ordered levels, lowest first the expected level, the distribution, a confidence

Confidence is (N * p_max - 1) / (N - 1) for a choice: 0 for a uniform answer, 1 for a certain one, whatever the option count. For a score it is based on the spread over the levels, so mass on two neighbouring levels still counts as confident.

🧠 How it works

The language-model head of the base model is dropped. A pointer head of about a million parameters takes the hidden state at the final <answer> token as a query and the hidden state at the last token of each option's line as keys, and scores one against the others. No weight belongs to an option slot, so option lists are defined by the request and the model cannot learn that the first option is usually right. Option order is shuffled on every training row.

<state>
Help! My payouts have been failing for 3 days!
</state>
<question type="choice">
Select exactly one option.
Which team should handle this?
<options>
1. billing
2. sales
3. retail
</options>
</question>
<answer>

⚖️ Buridan and Strands Decider

Same base model, prompt format, pointer head, LoRA setup and checkpoint layout as StrandsAgents/strands-decider-2B-hobson-v19 (v19), so buridan.py loads either one. The training recipe is a cut-down version of upstream's v13 recipe, which is v7's recipe on the Qwen3.5 torso:

Buridan Strands Decider v19
Base Qwen/Qwen3.5-2B-Base Qwen/Qwen3.5-2B-Base
Head pointer, dim 256, fp32 pointer, dim 256, fp32
Adapter LoRA r=16, alpha=32, attention, MLP and Gated DeltaNet projections same
Short classification tasks 17 21 (adds Yelp and three RuleTaker depths)
Multi-step documents (ContractNLI, MuSiQue, BoardgameQA) no 12,909 rows
Generated workplace questions and answer adequacy no 9,981 rows
Loss cross-entropy, ordinal smoothing 0.1, KL 0.3 to the frozen torso same, plus KL 1.0 to a parent model's answers on multi-step rows
Temperatures fitted per question type no yes
Training one run, one epoch parent model (about 5 h) then the final run (about 6 h) on an RTX 3090

🏋️ Training

Base Qwen/Qwen3.5-2B-Base
Data 82,449 rows from 17 tasks: ag_news, banking77, clinc150, dbpedia, lang_id, yahoo_topics, spam, toxicity, mnli_entail, boolq, paws, vitaminc, wnli, pubmed_qa, sst5_sentiment, app_reviews, formality
Steps 2,577 (one epoch), batch size 32
Learning rate 1e-4 for LoRA, 1e-3 for the head, cosine, 3% warmup
Mean train loss 0.5431
Hardware NVIDIA A10G, 151 min

📊 Results

Four tasks that neither model trained on, as in upstream's held-out evaluation: emotion, massive_intent, sarcasm, hate_severity. One seeded sample, the same rows for every model, scored by buridan.py eval. Upstream's model uses the temperatures stored with it; Buridan has none. ECE is the expected calibration error over 10 confidence bins (lower is better).

Model n Accuracy ECE Yes/no Choice Score
jgalego/buridan-2b 6000 0.631 0.088 0.585 0.725 0.434
StrandsAgents/strands-decider-2B-hobson-v19 6000 0.634 0.052 0.58 0.725 0.454

Accuracy per task:

Model emotion massive_intent sarcasm hate_severity
jgalego/buridan-2b 0.587 0.869 0.585 0.434
StrandsAgents/strands-decider-2B-hobson-v19 0.587 0.866 0.58 0.454

Share of answers and accuracy by confidence band. Upstream's convention is to act on answers at 0.9 or above, confirm those from 0.5 to 0.9 and send the rest to a person.

Model >= 0.9 0.5 - 0.9 < 0.5
jgalego/buridan-2b 15% at 0.961 30% at 0.755 54% at 0.468
StrandsAgents/strands-decider-2B-hobson-v19 23% at 0.947 31% at 0.668 46% at 0.457

For reference, upstream reports 0.653 (v7) and 0.641 (v19) on its own 6,000-row sample of these tasks. Upstream's main benchmark, JevBench, is not run here.

⚠️ Limitations

  • Trained on short classification only. Long, multi-step documents are upstream's weak spot even with its extra data, and this model has none of that data.
  • Upstream finds that on tasks unlike anything in training, choice questions are usable behind a confidence threshold and score questions are not.
  • Confidence is measured on short held-out classification. Check it on your own traffic before you set a threshold.
  • English only.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jgalego/buridan-2b

Adapter
(23)
this model

Datasets used to train jgalego/buridan-2b

Space using jgalego/buridan-2b 1

Collection including jgalego/buridan-2b