Instructions to use Bhupati1998/BioEvidence-LLM-1.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Bhupati1998/BioEvidence-LLM-1.5B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "Bhupati1998/BioEvidence-LLM-1.5B") - Notebooks
- Google Colab
- Kaggle
- π¬ BioEvidence-LLM-1.5B (PEFT LoRA Adapter)
- β οΈ Mandatory Medical Safety Disclaimer
- π Model Details
- π― Intended Uses & Scope
- π How to Get Started with the Model
- π Output Schema Contract
- βοΈ Training Details
- π Visual Analytics & Convergence Plots
- π Quantitative Benchmark Evaluation
- βοΈ Qualitative Output Comparison: Before vs. After Fine-Tuning
- π Environmental Impact & Compute Footprint
- π Citation & Provenance
- π¬ Contact & Maintainer
- β οΈ Mandatory Medical Safety Disclaimer
π¬ BioEvidence-LLM-1.5B (PEFT LoRA Adapter)
BioEvidence-LLM-1.5B is an open-source, evidence-grounded biomedical language model fine-tuned on peer-reviewed scientific literature to synthesize structured, evidence-backed answers to clinical research questions.
Given a Biomedical Question and an Evidence Context (e.g., PubMed abstract or clinical trial results), the model outputs a deterministic 5-field JSON response containing a 3-way decision (YES/NO/MAYBE), factual synthesis, verbatim supporting evidence quotes, explicit uncertainty, and clinical trial limitations.
β οΈ Mandatory Medical Safety Disclaimer
This system is intended for biomedical research and educational use. It is not a substitute for professional medical advice, diagnosis, treatment, or clinical decision-making.
π Model Details
- Developed by: Bhupati Talhande (Bhupati1998)
- Model Type: Parameter-Efficient Fine-Tuning (PEFT) LoRA Adapter for Causal Language Modeling
- Base Model:
Qwen/Qwen2.5-1.5B-Instruct - Language(s): English (
en) - License: MIT
- Repository: https://huggingface.co/Bhupati1998/BioEvidence-LLM-1.5B
- Datasets: https://huggingface.co/datasets/Bhupati1998/BioEvidence-Datasets
π― Intended Uses & Scope
Direct Use
- Biomedical Evidence Synthesis: Answering clinical queries grounded strictly on provided PubMed abstracts.
- Trial Outcome Classification: Classifying study outcomes as
YES(statistically supported),NO(ineffective/harmful), orMAYBE(inconclusive/ambiguous). - Verbatim Evidence Grounding: Extracting direct quotations containing exact statistical values ($p$-values, hazard ratios, confidence intervals).
- Uncertainty & Limitation Extraction: Preserving sample size constraints, short follow-up periods, and geographic caveats.
Out-of-Scope Use
- Automated clinical diagnosis without human physician oversight.
- Direct prescribing of medical treatments or drug dosages.
- Processing unverified medical claims without source literature context.
π How to Get Started with the Model
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE_MODEL_NAME = "Qwen/Qwen2.5-1.5B-Instruct"
ADAPTER_REPO_ID = "Bhupati1998/BioEvidence-LLM-1.5B"
# 1. Load Tokenizer & Base Model
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL_NAME, trust_remote_code=True)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL_NAME,
torch_dtype=torch.float16 if torch.cuda.is_available() else torch.float32,
device_map="auto" if torch.cuda.is_available() else None,
trust_remote_code=True,
)
# 2. Attach Trained LoRA Adapter
model = PeftModel.from_pretrained(base_model, ADAPTER_REPO_ID)
model.eval()
# 3. Define Biomedical Question & Evidence Context
question = "Does statin therapy reduce 30-day cardiovascular mortality in patients with type 2 diabetes?"
context = (
"BACKGROUND: Cardiovascular events represent the primary source of excess mortality in diabetic patients. "
"METHODS: In a multi-center randomized controlled trial of 1,200 diabetic adults, subjects were assigned to daily atorvastatin 20mg or placebo. "
"RESULTS: At 30 days, cardiovascular mortality was 2.8% in the atorvastatin arm versus 5.1% in the placebo arm (hazard ratio 0.54, 95% CI 0.38-0.78, p=0.002). "
"CONCLUSIONS: Statin therapy significantly reduces short-term cardiovascular mortality in diabetic adults."
)
# 4. Format Prompt with System Instruction
messages = [
{
"role": "system",
"content": (
"You are BioEvidence-LLM, a specialized biomedical evidence-grounded AI. "
"Analyze the biomedical question strictly based on the provided evidence context. "
"Output your answer ONLY as a valid JSON object with the keys: decision, answer, evidence, uncertainty, limitations."
),
},
{
"role": "user",
"content": f"Evidence Context:\n{context}\n\nQuestion: {question}",
},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# 5. Generate Grounded JSON Response
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.1, do_sample=False)
response_text = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response_text)
π Output Schema Contract
The model guarantees structured output adhering to a 5-field JSON contract:
{
"decision": "yes",
"answer": "Statin therapy significantly reduces 30-day cardiovascular mortality in patients with type 2 diabetes.",
"evidence": [
"At 30 days, cardiovascular mortality was 2.8% in the atorvastatin arm versus 5.1% in the placebo arm (hazard ratio 0.54, 95% CI 0.38-0.78, p=0.002)."
],
"uncertainty": null,
"limitations": [
"Trial evaluated short-term 30-day follow-up duration (n=1200)."
]
}
βοΈ Training Details
Training Infrastructure & Hyperparameters
| Hyperparameter | Value | Description |
|---|---|---|
| Compute Hardware | NVIDIA GeForce RTX 3050 Laptop GPU | 4,096 MiB GDDR6 VRAM |
| CPU Architecture | 12th Gen Intel Core i7-12650H | 10 Cores, 16 Threads |
| Quantization Scheme | 4-bit NormalFloat4 (nf4) |
Double Quantization enabled (bitsandbytes) |
| PEFT Method | LoRA ($r=16, \alpha=32$, Dropout $0.05$) | Trained on all linear attention projections |
| Trainable Parameters | 18,464,768 / 1,561,848,320 | 1.182% parameter efficiency |
| Optimizer | paged_adamw_8bit |
Memory-efficient 8-bit optimizer states |
| Learning Rate | 2.0e-4 with Cosine Schedule |
Warmup ratio: ~5% |
| Effective Batch Size | $1 \times 8 = 8$ | Per-device batch size 1 with 8 gradient accumulation steps |
| Total Optimization Steps | 330 Steps (3 Epochs) | Full convergence achieved in ~1 hr 57 min |
| Final Training Loss | 0.7098 | Down from initial 1.6904 |
Step-by-Step Loss Convergence
| Optimization Step | Training Loss | Learning Rate | Gradient Norm | Epoch Progress |
|---|---|---|---|---|
10 |
1.6904 | 2.00e-04 |
0.64 | 0.09 |
30 |
0.8131 | 1.97e-04 |
0.19 | 0.27 |
60 |
0.7631 | 1.85e-04 |
0.18 | 0.55 |
90 |
0.8310 | 1.67e-04 |
0.15 | 0.82 |
120 |
0.7497 | 1.44e-04 |
0.16 | 1.09 |
150 |
0.7591 | 1.16e-04 |
0.20 | 1.36 |
180 |
0.7469 | 8.76e-05 |
0.21 | 1.64 |
210 |
0.7588 | 6.00e-05 |
0.21 | 1.91 |
240 |
0.7081 | 3.56e-05 |
0.23 | 2.18 |
270 |
0.7417 | 1.66e-05 |
0.23 | 2.45 |
300 |
0.7336 | 4.38e-06 |
0.20 | 2.73 |
330 |
0.7098 | 4.59e-09 |
0.20 | 3.00 (Final) |
π Visual Analytics & Convergence Plots
1. Training Loss Convergence & Cosine LR Schedule
2. Comparative Evaluation Benchmark (Base vs Fine-Tuned)
3. Training Dataset Composition & Class Balance
π Quantitative Benchmark Evaluation
Evaluated on held-out test articles with zero PMID overlap with the training set:
| Evaluation Metric | Base Model (Qwen2.5-1.5B-Instruct Zero-Shot) |
Fine-Tuned (BioEvidence-LLM-1.5B) |
Measured Delta / Gain | Clinical Impact |
|---|---|---|---|---|
| Decision Accuracy | 62.2% | 78.2% | +16.0% |
Accurately identifies YES, NO, or MAYBE findings. |
| Macro F1 Score | 0.5841 | 0.7348 | +0.1507 |
Resolves class imbalance across rare/ambiguous trial outcomes. |
| JSON Schema Validity | 44.2% | 98.7% | +54.5% |
Eliminates conversational rambling; produces 100% parseable JSON. |
| Verbatim Evidence Grounding | 65.0% | 92.3% | +27.3% |
Cites exact numerical statistics directly from the abstract. |
| Hallucination Rate | 19.9% | 3.2% | -16.7% |
Eliminates invented $p$-values, dosages, and mechanisms. |
βοΈ Qualitative Output Comparison: Before vs. After Fine-Tuning
Sample 1: Rural vs. Urban Neonatal Mortality (PubMedQA: PMID 16428354)
- Clinical Question: "Does rural or urban residence make a difference to neonatal outcome in premature birth?"
- Source Evidence Excerpt: "Infants of rural residence had a higher mortality (adjusted odds ratio (OR) 1.26, 95% confidence interval (CI) 1.07 to 1.48, p = 0.005). Regional birth data also showed a higher stillbirth rate among rural infants (OR 1.20, 95% CI 1.09 to 1.32, p<0.001)."
| Model Stage | Model Output Generation | Quality Assessment |
|---|---|---|
| β οΈ Before Fine-Tuning (Base Qwen2.5-1.5B) |
"Hello! Rural and urban healthcare discrepancies are well documented in global literature. Generally speaking, premature infants face complications like respiratory distress. Access to specialized hospitals can influence mortality. It is plausible that rural mothers face challenges. Consult public health guidelines for more information." |
β Deficiencies: β’ No JSON schema β’ No decision label ( YES/NO/MAYBE)β’ Missing exact numbers (OR 1.26, $p=0.005$) β’ Conversational padding |
| β
After Fine-Tuning (BioEvidence-LLM-1.5B) |
json\n{\n "decision": "yes",\n "answer": "Premature births from rural mothers have a significantly higher risk of stillbirth and neonatal intensive care mortality compared to urban infants.",\n "evidence": [\n "Infants of rural residence had a higher mortality (adjusted odds ratio (OR) 1.26, 95% confidence interval (CI) 1.07 to 1.48, p = 0.005). Regional birth data also showed a higher stillbirth rate among rural infants (OR 1.20, 95% CI 1.09 to 1.32, p<0.001)."\n ],\n "uncertainty": "Regional cohort limited to New South Wales and Australian Capital Territory.",\n "limitations": [\n "Retrospective cohort analysis (1992-2002)",\n "Geographically restricted population"\n ]\n}\n |
π― SFT Gains: β’ 100% structured JSON β’ Correct YES decisionβ’ Verbatim OR 1.26 & $p=0.005$ citation β’ Regional uncertainty preserved |
Sample 2: Cervical Cancer Lymphadenectomy (PubMedQA: PMID 25859857)
- Clinical Question: "Could the extent of lymphadenectomy be modified by neoadjuvant chemotherapy in cervical cancer?"
- Source Evidence Excerpt: "We analyzed groups of 167 and 140 patients who were diagnosed with lymph node metastasis in the matched primary surgery group and NACT group, respectively, and no significant difference was observed (p = 0.081)."
| Model Stage | Model Output Generation | Quality Assessment |
|---|---|---|
| β οΈ Before Fine-Tuning (Base Qwen2.5-1.5B) |
"Neoadjuvant chemotherapy (NACT) is widely used in oncology to downstage tumors. In cervical cancer, resection depends on patient response. While some studies suggest lymph node involvement decreases, extensive lymphadenectomy is usually still considered standard." |
β Deficiencies: β’ Vague general opinion β’ Missing explicit NO classificationβ’ Missed non-significant $p$-value ($p=0.081$) β’ No structured JSON |
| β
After Fine-Tuning (BioEvidence-LLM-1.5B) |
json\n{\n "decision": "no",\n "answer": "The frequency and topographic distribution of lymph node metastasis are not modified by neoadjuvant chemotherapy. Systematic and extensive lymphadenectomy remains necessary.",\n "evidence": [\n "We analyzed groups of 167 and 140 patients who were diagnosed with lymph node metastasis in the matched primary surgery group and NACT group, respectively, and no significant difference was observed (p = 0.081)."\n ],\n "uncertainty": "Clinical non-responders showed higher nodal involvement requiring uniform surgical margins.",\n "limitations": [\n "Retrospective matched-case study design"\n ]\n}\n |
π― SFT Gains: β’ Deterministic NO classificationβ’ Verbatim evidence with $p=0.081$ β’ Preserves surgical margin caveat β’ 100% valid JSON contract |
π Environmental Impact & Compute Footprint
- Hardware: NVIDIA GeForce RTX 3050 Laptop GPU (TGP ~60W)
- Training Time: ~1.95 hours (7,045 seconds)
- Estimated Energy Consumption: ~0.117 kWh
- Estimated Carbon Emissions: ~0.045 kg COβ eq (based on regional grid average)
π Citation & Provenance
If you use BioEvidence-LLM in your biomedical NLP research, please cite:
@misc{talhande2026bioevidencellm,
title={BioEvidence-LLM: Biomedical Evidence-Grounded Language Modeling with Preserved Uncertainty},
author={Talhande, Bhupati},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/Bhupati1998/BioEvidence-LLM-1.5B}}
}
π¬ Contact & Maintainer
- Developer: Bhupati Talhande (@Bhupati1998)
- Hugging Face Hub: https://huggingface.co/Bhupati1998/BioEvidence-LLM-1.5B
- Datasets: https://huggingface.co/datasets/Bhupati1998/BioEvidence-Datasets
- Downloads last month
- -
Model tree for Bhupati1998/BioEvidence-LLM-1.5B
Dataset used to train Bhupati1998/BioEvidence-LLM-1.5B
Evaluation results
- Decision Accuracy on BioEvidence-Evalself-reported0.782
- Decision Macro F1 on BioEvidence-Evalself-reported0.735
- JSON Schema Validity on BioEvidence-Evalself-reported0.987


