Instructions to use EPFLiGHT/EuroLLM-22B-MeditronFO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EPFLiGHT/EuroLLM-22B-MeditronFO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="EPFLiGHT/EuroLLM-22B-MeditronFO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("EPFLiGHT/EuroLLM-22B-MeditronFO") model = AutoModelForCausalLM.from_pretrained("EPFLiGHT/EuroLLM-22B-MeditronFO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use EPFLiGHT/EuroLLM-22B-MeditronFO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EPFLiGHT/EuroLLM-22B-MeditronFO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EPFLiGHT/EuroLLM-22B-MeditronFO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/EPFLiGHT/EuroLLM-22B-MeditronFO
- SGLang
How to use EPFLiGHT/EuroLLM-22B-MeditronFO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "EPFLiGHT/EuroLLM-22B-MeditronFO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EPFLiGHT/EuroLLM-22B-MeditronFO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "EPFLiGHT/EuroLLM-22B-MeditronFO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EPFLiGHT/EuroLLM-22B-MeditronFO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use EPFLiGHT/EuroLLM-22B-MeditronFO with Docker Model Runner:
docker model run hf.co/EPFLiGHT/EuroLLM-22B-MeditronFO
EuroLLM-22B-MeditronFO
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
We're introducing EuroLLM-22B-MeditronFO, our latest fully open medical specialist LLM, medical specialization of EuroLLM-22B-Instruct on the Fully Open Meditron Corpus. This model is part of the Fully Open Meditron family — the first end-to-end auditable pipeline for clinical LLMs, with open weights, open data, open training recipe, and clinician-vetted corpus construction.
- Part of the Fully Open Meditron family: End to end fully open clinical LLMs
- Preferred over EuroLLM-22B-Instruct in 59.2% of comparisons on AutoMOOVE, the clinician-validated LLM-judge evaluation (three-judge majority, 516 clinician-written vignettes)
Benchmark
Accuracy (%) on four multiple-choice medical benchmarks (greedy decoding) and score (%) on HealthBench (all 5,000 conversations, Gemma-4-31B grader). See the paper for the full evaluation details, confidence intervals and the AutoMOOVE results.
| Benchmark | EuroLLM-22B-Instruct | EuroLLM-22B-MeditronFO | Δ |
|---|---|---|---|
| MedMCQA | 55.8 | 55.2 | −0.6 |
| MedQA | 65.3 | 64.7 | −0.6 |
| PubMedQA | 76.4 | 77.9 | +1.5 |
| MedXpertQA | 14.6 | 14.9 | +0.3 |
| HealthBench | 39.9 | 43.5 | +3.6 |
| Average (5) | 50.40 | 51.23 | +0.83 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "EPFLiGHT/EuroLLM-22B-MeditronFO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "user", "content": "A 62-year-old woman presents with a three-day history of dyspnea on exertion and a productive cough. What is the differential diagnosis?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Training
- Base model: EuroLLM-22B-Instruct
- Corpus: Fully Open Meditron, 533,888 examples: Curated QA (195.5k, seven public medical QA datasets), Synthetic Curated QA (184.5k), Guidelines QA (129.8k, generated from 14,192 clinical practice guidelines) and Synthetic vignettes (24.0k, modelled on clinician-written MOOVE vignettes). Every answer is written by gpt-oss-120b with rejection sampling (pass@8; multiple-choice answers must match the gold label), with generation prompts co-authored by four clinicians
- Decontamination: two-stage n-gram decontamination against all evaluation benchmarks and the AutoMOOVE test set (4,197 rows removed)
- Recipe: 2 epochs, learning rate 1e-5 with cosine decay and 10% warmup, effective batch 128 sequences of 8,192 packed tokens, AdamW, seed 42
- Framework: Axolotl with FSDP v2, bf16
Full hyperparameters are in the training appendix of the paper.
Compute & footprint
The training was done on 8 nodes of 4 NVIDIA GH200 GPUs for 4 h 01 min (129 GPU-hours) on the Alps supercomputer of the CSCS Swiss National Supercomputing Centre. Our trainings have a carbon neutral footprint as the CSCS data center is carbon neutral (CSCS energy efficiency).
Limitations & intended use
MeditronFO can produce text on a variety of topics, but the generated content may not always be factually accurate, logically consistent, or free from biases present in the training data. MeditronFO has been trained to be specialised for Medicine and is intended to be used for Medicine related tasks evaluation. These models should be used as assistive tools rather than definitive sources of information. Users should always verify important information and critically evaluate any generated content.
Citation
If you find MeditronFO useful in your research, please cite our preprint:
@misc{theimerlienhard2026fullyopenmeditronauditable,
title = {Fully Open Meditron: An Auditable Pipeline for Clinical LLMs},
author = {Xavier Theimer-Lienhard and Mushtaha El-Amin and Fay Elhassan and Sahaj Vaidya and Victor Cartier-Negadi and David Sasu and Lars Klein and Mary-Anne Hartley},
year = {2026},
eprint = {2605.16215},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2605.16215}
}
Contact
Please use the community tab for any discussions or issue related to this model. Questions related to the project can be sent to xavier.theimer-lienhard@epfl.ch or mary-anne.hartley@epfl.ch.
- Downloads last month
- 50