Instructions to use xjc1022/MELP-Encoder-repro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xjc1022/MELP-Encoder-repro with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="xjc1022/MELP-Encoder-repro", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("xjc1022/MELP-Encoder-repro", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
MELP ECG encoder (reproduction)
The ECG tower of a MELP model trained from this fork:
HKU-MedAI/MELP#4. Same architecture and
packaging as fuyingw/MELP_Encoder, so
it is a drop-in replacement. Original paper: From Token to Rhythm: A Multi-Scale
Approach for ECG-Language Pretraining (ICML 2025),
Wang, Xu & Yu — all credit for the method belongs to the authors.
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("xjc1022/MELP-Encoder-repro", trust_remote_code=True).eval()
ecg = torch.rand(1, 12, 5000) # 12 leads, 10 s at 500 Hz, scaled to [0, 1]
with torch.no_grad():
out = model(ecg)
# out["proj_ecg_emb"] (B, 256) rhythm level, the zero-shot embedding
# out["ecg_beat_emb"] (B, 12, 256) beat level
# out["ecg_token_emb"] (B, 128, 768) token level
Lead order is I, II, III, aVR, aVF, aVL, V1-V6, and each record is min-max scaled to [0, 1] over the whole 12x5000 array — the same preprocessing the training data used.
Zero-shot results
Six-dataset mean AUROC from scripts/zeroshot/test_zeroshot.py (the paper's protocol),
test splits:
| Rhythm | Form | Sub | Super | CPSC2018 | CSN | Average | |
|---|---|---|---|---|---|---|---|
| This run | 86.43 | 70.32 | 77.65 | 77.22 | 84.18 | 74.58 | 78.40 |
| Paper (Table 3) | 85.4 | 69.1 | 81.2 | 76.2 | 84.2 | 77.6 | 79.0 |
Those numbers describe the full MELP model this encoder came from: zero-shot needs the text tower as well, which is not part of this repo. What is published here is the ECG tower alone, for use as a feature extractor.
How it was trained
Text tower fuyingw/heart_bert, ECG tower initialised from an ECG-FM wav2vec2-CMSC
checkpoint, trained on MIMIC-IV-ECG report pairs. 4x RTX 3090, batch 64/device,
lr 1e-4, n_queries_contrast=12, loss weights 1.0 / 2.0 / 0.2. Best checkpoint by
validation zero-shot AUROC, epoch 4.
Three things mattered more than any hyperparameter here, and all three are fixes in the linked PR:
- The shipped LR scheduler has no warmup, and without it both towers collapse within ~25 steps — every zero-shot AUROC comes out at exactly 0.5.
- Training on denoised waveforms while evaluating on raw ones opens a domain gap worth 5.6 AUROC points (70.75 -> 78.40 once the raw wfdb records are used).
- The peak is at epoch 4. Runs stopped earlier, or validated only at epoch boundaries, miss it.
Data
Pretrained on MIMIC-IV-ECG, which is PhysioNet credentialed-access data. These are model weights rather than data, but check the PhysioNet data use agreement before redistributing anything derived from them.
- Downloads last month
- 24