S1-PubMedBERT: a calibrated single-pass decision model for biomedical text

S1-PubMedBERT answers typed questions about a text in one forward pass and returns temperature-scaled probabilities that can be thresholded directly. It is a System One decision model in the sense of Jev and Laya, built with the sbert2s1 framework by converting the retrieval encoder S-PubMedBert-MS-MARCO into a cross mask-slot decision model (the "C" architecture).

Three question types share one request schema (the public Jev/Laya schema):

Type You give You get
choice a dict of option keys and descriptions the chosen key, a probability per option, the confidence
score an ordered list of rubric levels (low to high) the expected level, a probability per level
noul a statement-style instruction P(true)

Several questions about the same text can be asked in one call; outputs cannot violate the schema.

Usage

pip install torch transformers safetensors huggingface_hub
import importlib.util
from huggingface_hub import hf_hub_download

# the inference code ships with the model (read it first: about 300 lines of plain PyTorch)
path = hf_hub_download("pritamdeka/S1-PubMedBERT", "sbert2s1.py")
spec = importlib.util.spec_from_file_location("sbert2s1", path)
sbert2s1 = importlib.util.module_from_spec(spec)
spec.loader.exec_module(sbert2s1)

model = sbert2s1.load("pritamdeka/S1-PubMedBERT")        # GPU if available, otherwise CPU

abstract = ("In this double-blind trial, 4,012 adults with atrial fibrillation were randomised to drug X "
            "or placebo. Drug X reduced stroke by 29% (HR 0.71, 95% CI 0.60-0.84); major bleeding was "
            "more frequent with drug X (3.1% vs 2.2%).")

questions = {
    "design": {"type": "choice", "instructions": "What is the study design?",
               "criteria": {"rct": "randomised controlled trial", "cohort": "cohort study",
                            "case_control": "case-control study", "case_report": "case report",
                            "review": "review or meta-analysis", "other": "other design"}},
    "humans": {"type": "noul", "instructions": "Does the study involve human participants?"},
    "adverse": {"type": "noul", "instructions": "Does the abstract report adverse effects of a drug?"},
    "evidence": {"type": "score", "instructions": "How strong is the evidence that drug X prevents stroke?",
                 "criteria": ["none", "weak", "moderate", "strong"]},
}

print(model.predict(abstract, questions))

Output format (illustrative placeholders, not real predictions):

{"answers": {
  "design":   {"type": "choice", "choice": "<key>", "answer_confidence": <p_max>,
               "probabilities": {"rct": <p>, "cohort": <p>, ...}},
  "humans":   {"type": "noul", "noul": <P(true)>, "answer_confidence": <max(P(true), P(false))>},
  "adverse":  {"type": "noul", "noul": <P(true)>, "answer_confidence": <...>},
  "evidence": {"type": "score", "score": <expected level 0..3>, "answer_confidence": <p_max>,
               "probabilities": {"0": <p>, "1": <p>, "2": <p>, "3": <p>}}}}

Batches. model.predict_batch([(text1, questions1), (text2, questions2), ...], batch_size=32). Raw logits. model.logits([...]) returns the unscaled logits per question. Using the confidence. answer_confidence (or noul) is the quantity to threshold, e.g. accept automatically above 0.9 and send the rest to a person or a larger model.

Inputs. The state may be a string or a JSON-serialisable record (it is serialised to text). noul questions may define criteria: {"false": ..., "true": ...} and optional labels. The model reads at most 512 tokens: the question block (instructions plus options, at most 192 tokens) and then the text. Longer texts are truncated from the end, and the answer is marked "truncated": true; for long documents, split the text and ask per chunk.

The laya/ folder contains the same weights in the Laya checkpoint layout for Laya-compatible runtimes.

Training

  • Initialisation: pritamdeka/S-PubMedBert-MS-MARCO (PubMedBERT further trained for MS MARCO retrieval), 110M parameters, plus a two-layer mask-slot decision head (Laya-compatible).
  • Objective: cross-entropy against the gold option distribution. In our experiments this matched or beat the released RLCD reinforcement-learning recipe (70.2 vs 67.7 seen-task points, 3 seeds each).
  • Data (public only; no credentialed clinical data): PubMedQA (labelled and artificial), SciFact, HealthVer, DDI-2013, Hallmarks of Cancer, ADE, Druglib reviews, MedQA, MedMCQA, and MEDLINE-S1 (243k decisions derived from NLM MEDLINE indexing: study design, disease area, human/animal subjects, adverse effects, drug therapy, age/sex groups). Options are shuffled during training (except ordinal rubrics).
  • Recipe: 8,000 steps, batch 32, AdamW (encoder 2.5e-5, head 1e-4), cosine schedule, bf16, temperature-sampled task mixture.
  • Calibration: one temperature per (question type, option-count bucket), fitted on calibration slices carved from the training data before training.
  • Selection: three seeds were trained; this checkpoint is the seed with the best seen-task accuracy (the 3-seed mean is 70.2 ± 0.4, so the number below is slightly optimistic).

Evaluation

Chance-normalised accuracy (acc - 1/K)/(1 - 1/K) (0 = chance, 100 = perfect), averaged over question groups within a task; ECE after temperature scaling (15 bins). Test splits; x100.

Task (seen in training) acc_cn ECE
PubMedQA 47.8 4.3
SciFact 61.2 8.8
HealthVer 56.0 2.8
DDI-2013 92.9 4.1
Hallmarks of Cancer (10 noul questions) 88.8 2.4
ADE 88.6 1.3
Druglib (3 score questions) 41.6 4.3
MEDLINE-S1 87.1 2.4
Mean of the 8 seen tasks 70.5 3.8
Not trained on acc_cn ECE
PUBHEALTH (claim veracity) 22.7 13.1
MTSamples (specialty routing) 72.2 27.5
BIOSSES (similarity rubric) -5.0 16.2
Exam questions (MedQA / MedMCQA / MMLU-medical) 14.8 / 13.1 / 14.8 1.3-3.0
General-domain typed-decisions 8.7 16.5

Intended use and limitations

  • Intended use: research and prototyping of fast, calibrated decisions on biomedical literature and similar text, close to the training tasks (evidence classification, claim verification, drug interactions/adverse events, indexing-style questions), typically as the first stage of a pipeline that escalates low-confidence cases.
  • Not a medical device and not validated for clinical decisions. It was not trained on clinical notes and is weak on them. Calibration was measured on in-distribution test sets; on new tasks the probabilities can be badly miscalibrated (see MTSamples, PUBHEALTH). Re-fit temperatures (or validate thresholds) on your own labelled sample before relying on the confidence.
  • Out-of-domain and knowledge questions: near chance on general-domain decisions and on exam questions that require memorised medical knowledge; a 110M encoder does not store that knowledge.
  • Sensitivity to wording: answers depend on how instructions and options are phrased, and the cross head is somewhat sensitive to option order (about 3% of seen-task answers change under option permutation).
  • Licence: CC BY-NC 2.0, inherited from the base model; non-commercial use only. Check the licences of the training datasets for your use case (several are themselves non-commercial).

Citation

@misc{deka2026retrievaltypeddecisionscalibrated,
      title={From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders}, 
      author={Pritam Deka},
      year={2026},
      eprint={2610.02486},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.02486}, 
}
@article{deka2022improved,
  title   = {Improved Methods To Aid Unsupervised Evidence-Based Fact Checking For Online Health News},
  author  = {Deka, Pritam and Jurek-Loughrey, Anna and Deepak, P.},
  journal = {Journal of Data Intelligence}, volume = {3}, number = {4}, pages = {474--504}, year = {2022}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pritamdeka/S1-PubMedBERT

Finetuned
(5)
this model

Paper for pritamdeka/S1-PubMedBERT