BioJev-9B / README.md
Gabriel382's picture
BioJev_9B_HF_Supplement_v1
23f64d4
|
Raw History Blame Contribute Delete
5.16 kB
---
library_name: peft
base_model: Qwen/Qwen3.5-9B-Base
pipeline_tag: text-classification
tags:
- biojev
- biomedical
- qwen3.5
- peft
- qlora
- natural-language-inference
- biomedical-nlp
- system-one
language:
- en
---
# BioJev-9B
**BioJev-9B** is the 9B member of the BioJev biomedical decision-model family.
Built from **Qwen/Qwen3.5-9B-Base**, it follows the full BioJev pipeline:
```text
Qwen3.5-9B-Base
↓
Biomedical DAPT — PubMed + PMC
↓
General NLI — SNLI + MNLI + ANLI
↓
Biomedical NLI — BioNLI + NLI4CT
↓
BioJev-9B
```
The released checkpoint is a **PEFT/QLoRA sequence-classification adapter** with:
```text
0 → contradiction
1 → entailment
2 → neutral
```
## Links
- BioJev-9B: https://huggingface.co/Gabriel382/BioJev-9B
- BioJev-4B: https://huggingface.co/Gabriel382/BioJev
- BioJev-Nano: https://huggingface.co/Gabriel382/BioJev-Nano
- Source code: https://github.com/Gabriel382/BioJev
## Authors
**Creator:** Gabriel Henrique Alencar Medeiros
**Supervisor:** Lina F. Soualmia
Developed at **LITIS / Université de Rouen Normandie**.
## Training recipe
### Biomedical DAPT
```text
100M biomedical tokens
80% PubMed / 20% PMC
sequence length 2048
QLoRA
bfloat16
4-bit quantization
```
### General NLI
```text
SNLI 35,000
MNLI 45,000
ANLI R1 20,000
----------------
Total 100,000
```
Dev results:
```text
Accuracy: 89.20%
Macro-F1: 89.15%
Loss: 0.2958
```
### Biomedical NLI
```text
BioNLI 30,000
NLI4CT 10,000
----------------
Total 40,000
```
Dev results:
```text
Accuracy: 92.77%
Macro-F1: 92.30%
Loss: 0.1996
```
## Frozen evaluation
| Dataset | Role | Macro-F1 |
|---|---|---:|
| BioNLI | held-out biomedical NLI | **94.73** |
| NLI4CT | validation seen during full training | **72.50** |
| ChemProt | zero-shot relation typing | **31.13** |
| DDI2013 | zero-shot relation typing | **36.35** |
| BioRED | zero-shot relation typing | **34.32** |
**Zero-shot relation-transfer mean:** **33.93 macro-F1**
**All-5 mean:** **53.80 macro-F1**
### Reliability
| Dataset | Accuracy (%) | Macro-F1 (%) | ECE (%) | NLL | Mean confidence (%) |
|---|---:|---:|---:|---:|---:|
| BioNLI | 95.14 | 94.73 | 2.20 | 0.146 | 97.27 |
| NLI4CT | 72.50 | 72.50 | 14.70 | 0.693 | 86.30 |
| ChemProt | 32.24 | 31.13 | 10.34 | 1.995 | 41.00 |
| DDI2013 | 39.33 | 36.35 | 9.17 | 1.372 | 32.27 |
| BioRED | 50.11 | 34.32 | 6.52 | 1.156 | 54.64 |
### Scaling context
```text
Relation-transfer mean macro-F1
BioJev-Nano 21.01
BioJev-4B 30.83
BioJev-9B 33.93
```
The 9B model improves aggregate transfer relative to 4B, with strong gains on ChemProt and BioRED, while DDI2013 does not improve monotonically with model size.
This is an empirical model-size comparison, not a formal scaling-law study.
# Loading BioJev-9B
Install:
```bash
pip install -U torch transformers peft accelerate
```
Optional 4-bit inference:
```bash
pip install -U bitsandbytes
```
Because BioJev is a three-class PEFT sequence classifier, reconstruct the base model with `num_labels=3` before loading the adapter.
```python
import torch
from peft import PeftConfig, PeftModelForSequenceClassification
from transformers import AutoModelForSequenceClassification, AutoTokenizer
MODEL = "Gabriel382/BioJev-9B"
LABEL2ID = {"contradiction": 0, "entailment": 1, "neutral": 2}
ID2LABEL = {v: k for k, v in LABEL2ID.items()}
cfg = PeftConfig.from_pretrained(MODEL)
base = AutoModelForSequenceClassification.from_pretrained(
cfg.base_model_name_or_path,
num_labels=3,
label2id=LABEL2ID,
id2label=ID2LABEL,
dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
device_map="auto" if torch.cuda.is_available() else None,
)
model = PeftModelForSequenceClassification.from_pretrained(
base, MODEL, is_trainable=False
)
tokenizer = AutoTokenizer.from_pretrained(MODEL)
if tokenizer.pad_token_id is None:
tokenizer.pad_token = tokenizer.eos_token
model.config.pad_token_id = tokenizer.pad_token_id
model.eval()
```
# System One compatibility API
The main BioJev repository exposes BioJev checkpoints through:
```text
POST /v1/systemone
```
Supported types:
```text
choice
noul
score
```
This is an engineering compatibility bridge over BioJev's NLI classifier, not a native Ollama/System-One checkpoint.
Serve a local copy with:
```bash
python scripts/serve_systemone.py \
--checkpoint /path/to/BioJev-9B \
--model-name biojev-9b \
--load-in-4bit \
--host 0.0.0.0 \
--port 8000
```
## Scientific status
BioJev is a research project. The model and System One compatibility layer are **not validated clinical decision systems** and must not be used as the sole basis for diagnosis, treatment, or other high-stakes medical decisions.
## Citation
A formal paper citation will be added when the BioJev publication is available. Until then, please cite the BioJev repository and the specific Hugging Face checkpoint used.
## Acknowledgements
BioJev-9B was created by **Gabriel Henrique Alencar Medeiros** under the supervision of **Lina F. Soualmia** at **LITIS / Université de Rouen Normandie**.