Instructions to use jgalego/buridan-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jgalego/buridan-2b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
buridan-2b
A decision model: it picks one of the options it is given and says how sure it is, in one forward pass, without generating text. A replication of Strands Decider on Qwen/Qwen3.5-2B-Base. buridan.py in this repo trains, evaluates and loads it.
Named after Buridan's ass, the donkey that starved between two equal piles of hay because it could not choose. This one always chooses. When the piles look equal, its confidence says so.
🚀 Usage
uv run https://huggingface.co/jgalego/buridan-2b/resolve/main/buridan.py ask \
--state "Help! My payouts have been failing for 3 days!" \
--choice "Which team should handle this?=billing,sales,retail" \
--noul "Does this convey urgency?" \
--score "How frustrated is the writer?=calm,frustrated,depressed"
Three question types share one head:
| Type | Options | Returns |
|---|---|---|
noul |
false, true | P(true) |
choice |
2 or more, set per request | the best option, every option's probability, a confidence |
score |
ordered levels, lowest first | the expected level, the distribution, a confidence |
Confidence is (N * p_max - 1) / (N - 1) for a choice: 0 for a uniform answer, 1 for a certain one, whatever the option count. For a score it is based on the spread over the levels, so mass on two neighbouring levels still counts as confident.
🧠 How it works
The language-model head of the base model is dropped. A pointer head of about a million parameters takes the hidden state at the final <answer> token as a query and the hidden state at the last token of each option's line as keys, and scores one against the others. No weight belongs to an option slot, so option lists are defined by the request and the model cannot learn that the first option is usually right. Option order is shuffled on every training row.
<state>
Help! My payouts have been failing for 3 days!
</state>
<question type="choice">
Select exactly one option.
Which team should handle this?
<options>
1. billing
2. sales
3. retail
</options>
</question>
<answer>
⚖️ Buridan and Strands Decider
Same base model, prompt format, pointer head, LoRA setup and checkpoint layout as StrandsAgents/strands-decider-2B-hobson-v19 (v19), so buridan.py loads either one. The training recipe is a cut-down version of upstream's v13 recipe, which is v7's recipe on the Qwen3.5 torso:
| Buridan | Strands Decider v19 | |
|---|---|---|
| Base | Qwen/Qwen3.5-2B-Base | Qwen/Qwen3.5-2B-Base |
| Head | pointer, dim 256, fp32 | pointer, dim 256, fp32 |
| Adapter | LoRA r=16, alpha=32, attention, MLP and Gated DeltaNet projections | same |
| Short classification tasks | 17 | 21 (adds Yelp and three RuleTaker depths) |
| Multi-step documents (ContractNLI, MuSiQue, BoardgameQA) | no | 12,909 rows |
| Generated workplace questions and answer adequacy | no | 9,981 rows |
| Loss | cross-entropy, ordinal smoothing 0.1, KL 0.3 to the frozen torso | same, plus KL 1.0 to a parent model's answers on multi-step rows |
| Temperatures fitted per question type | no | yes |
| Training | one run, one epoch | parent model (about 5 h) then the final run (about 6 h) on an RTX 3090 |
🏋️ Training
| Base | Qwen/Qwen3.5-2B-Base |
| Data | 82,449 rows from 17 tasks: ag_news, banking77, clinc150, dbpedia, lang_id, yahoo_topics, spam, toxicity, mnli_entail, boolq, paws, vitaminc, wnli, pubmed_qa, sst5_sentiment, app_reviews, formality |
| Steps | 2,577 (one epoch), batch size 32 |
| Learning rate | 1e-4 for LoRA, 1e-3 for the head, cosine, 3% warmup |
| Mean train loss | 0.5431 |
| Hardware | NVIDIA A10G, 151 min |
📊 Results
Four tasks that neither model trained on, as in upstream's held-out evaluation: emotion, massive_intent, sarcasm, hate_severity. One seeded sample, the same rows for every model, scored by buridan.py eval. Upstream's model uses the temperatures stored with it; Buridan has none. ECE is the expected calibration error over 10 confidence bins (lower is better).
| Model | n | Accuracy | ECE | Yes/no | Choice | Score |
|---|---|---|---|---|---|---|
| jgalego/buridan-2b | 6000 | 0.631 | 0.088 | 0.585 | 0.725 | 0.434 |
| StrandsAgents/strands-decider-2B-hobson-v19 | 6000 | 0.634 | 0.052 | 0.58 | 0.725 | 0.454 |
Accuracy per task:
| Model | emotion | massive_intent | sarcasm | hate_severity |
|---|---|---|---|---|
| jgalego/buridan-2b | 0.587 | 0.869 | 0.585 | 0.434 |
| StrandsAgents/strands-decider-2B-hobson-v19 | 0.587 | 0.866 | 0.58 | 0.454 |
Share of answers and accuracy by confidence band. Upstream's convention is to act on answers at 0.9 or above, confirm those from 0.5 to 0.9 and send the rest to a person.
| Model | >= 0.9 | 0.5 - 0.9 | < 0.5 |
|---|---|---|---|
| jgalego/buridan-2b | 15% at 0.961 | 30% at 0.755 | 54% at 0.468 |
| StrandsAgents/strands-decider-2B-hobson-v19 | 23% at 0.947 | 31% at 0.668 | 46% at 0.457 |
For reference, upstream reports 0.653 (v7) and 0.641 (v19) on its own 6,000-row sample of these tasks. Upstream's main benchmark, JevBench, is not run here.
⚠️ Limitations
- Trained on short classification only. Long, multi-step documents are upstream's weak spot even with its extra data, and this model has none of that data.
- Upstream finds that on tasks unlike anything in training, choice questions are usable behind a confidence threshold and score questions are not.
- Confidence is measured on short held-out classification. Check it on your own traffic before you set a threshold.
- English only.
- Downloads last month
- -
Model tree for jgalego/buridan-2b
Base model
Qwen/Qwen3.5-2B-Base