From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
Abstract
Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.
Community
This paper introduces SBERT2S1, a framework that adapts biomedical Sentence-Transformers into single-pass, schema-constrained "System One" decision models evaluated across BIODECIDE (a benchmark of grounded and clinical tasks) and trained on MEDLINE-S1 (243k decisions curated from NLM indexing). Across eleven encoders and six parent-retriever pairs, contrastive retrieval pre-training significantly improves zero-shot option matching and aids low-resource fine-tuning when paired with a novel Prior-Fused Residual (PFR) head—which achieves near-complete invariance to option ordering (0.3% flip rate vs. 15.7% for cross-encoder heads)—while standard cross-encoder heads deliver higher raw accuracy. Furthermore, theoretical and empirical analyses reveal that the popular RLCD training objective trails standard cross-entropy by 2.5–3.0 points due to an inflated reward normalization baseline, an issue resolved by a simple leave-one-out estimator, with no training objective out-calibrating temperature-scaled cross-entropy. Finally, cascading these fast ~110M base encoders with larger LLMs by escalating only the 20% least-confident decisions outperforms using either model alone, with models, code, and datasets openly released.
Get this paper in your agent:
hf papers read 2610.02486 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
pritamdeka/MEDLINE-S1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper