Luna / README.md
loopback-kr's picture
Update README.md
7c1b186 verified
|
Raw History Blame Contribute Delete
7.89 kB
---
language:
- en
license: mit
tags:
- medical-imaging
- ophthalmology
- retinal-fundus
- vision-language-model
- multimodal
- foundation-model
- vl-bert
- siglip2
- segmentation
- zero-shot-classification
pipeline_tag: image-feature-extraction
library_name: pytorch
---
# Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
**Luna** is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs).
Instead of pixel reconstruction, it is pre-trained by **predicting masked clinical tokens** (disease label, age, sex,
eye laterality) from the image, while a **UNETR decoder jointly supervises vessel / optic-disc segmentation**.
Architecturally it is a single-stream VL-BERT: a **SigLIP-2 ViT-Base/16 (512px) vision encoder**, a **BERT-Base
masked-LM initialized from PubMedBERT**, and a **UNETR head** for anatomy.
> Paper: N/A
> Code: https://github.com/loopback-kr/Luna
## Model Details
| Attribute | Value |
|---|---|
| **Architecture** | SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder |
| **Parameters** | 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse |
| **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) |
| **Vision encoder init.** | `google/siglip2-base-patch16-512` |
| **Text encoder init.** | `NeuML/pubmedbert-base-embeddings` tokenizer, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens |
| **Pre-training objective** | Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, *not* pixel-level MAE reconstruction |
| **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) |
| **License** | MIT |
Prompt templates used during pre-training: `"A CFP image of {CLS}"`, `"age is {CLS} years old"`,
`"gender is {CLS}"`, `"this is {CLS} direction eye"`.
## Files in this repository
This is the full `accelerate` training state at epoch 399:
| File | Contents |
|---|---|
| `model.safetensors` / `pytorch_model.bin` | Model weights (243,122,497 parameters, fp32) |
| `optimizer.bin`, `scheduler.bin`, `scaler.pt` | AdamW / LR-scheduler / GradScaler state, for exact resumption |
| `random_states_{0..3}.pkl` | RNG states of the four training processes |
## Usage
Full training, fine-tuning and evaluation code lives in the
[GitHub repository](https://github.com/loopback-kr/Luna) — see its `README.md` for the *Upstream training*,
*Downstream training*, *Zero-Shot* and *Segmentation* sections. Checkpoint paths there expect
`pytorch_model.bin`; passing the containing directory instead resumes optimizer and scheduler state as well.
## Training Data
**\*** = private institutional data, not available externally.
**Pre-training datasets**
| Dataset | Reference |
|---|---|
| LAG | Li, L. et al. Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model. *CVPR* (2019). |
| ODIR | Ocular Disease Intelligent Recognition (ODIR-2019). Peking University Grand Challenge (2019). |
| PARAGUAY | Castillo Benítez, V.E. et al. Dataset from fundus images for the study of diabetic retinopathy. *Data in Brief* 36, 107068 (2021). |
| G1020 | Bajwa, M.N. et al. G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma Detection. *IJCNN* (2020). |
| FUND-OCT | Hassan, T., Akram, M.U., Werghi, N. & Nazir, N. A hybrid convolutional framework for the automated extraction of retinal lesions and lesion-influenced grading of human retinal pathology. *IEEE J. Biomed. Health Inform.* 25(1) (2020). |
| Drishti-GS1 | Sivaswamy, J. et al. Drishti-GS: Retinal Image Dataset for Optic Nerve Head Segmentation. *ISBI* 53–56 (2014). |
| HRF | Budai, A. et al. Robust Vessel Segmentation in Fundus Images. *Int. J. Biomed. Imaging* 2013, 154860 (2013). |
| ORIGA | Zhang, Z. et al. ORIGA(-light): An Online Retinal Fundus Image Database for Glaucoma Analysis and Research. *EMBC* 3065–3068 (2010). |
| OIA-DDR | Li, T. et al. Diagnostic Assessment of Deep Learning Algorithms for Diabetic Retinopathy Screening. *Information Sciences* 501, 511–522 (2019). |
| SUSTech-SYSU | Lin, L. et al. The SUSTech-SYSU dataset for automated exudate detection and diabetic retinopathy grading. *Sci. Data* 7, 409 (2020). |
| JICHI | Takahashi, H. et al. Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy. *PLoS ONE* 12, e0179790 (2017). |
| CHAKSU | Kumar J H, R. et al. Chákṣu: A glaucoma specific fundus image database. *Sci. Data* 10, 70 (2023). |
| DR1 & DR2 | Pires, R. et al. Advancing Bag-of-Visual-Words Representations for Lesion Classification in Retinal Images. *PLoS ONE* 9, e96814 (2014). |
| DeepDRiD | Liu, R. et al. DeepDRiD: Diabetic Retinopathy—Grading and Image Quality Estimation Challenge. *Patterns* 3, 100512 (2022). |
| AIROGS-light-V2 | Kiefer, R. Glaucoma Dataset: EyePACS-AIROGS-light-V2. Kaggle (2024). doi:10.34740/KAGGLE/DSV/7802508 |
| FIDVS | Jin, K. et al. Fundus Image Dataset for Vessel Segmentation. Kaggle (2025). doi:10.34740/KAGGLE/DS/7319687 |
| JustRAIGS | Madadi, Y. et al. JustRAIGS: Justified Referral in AI Glaucoma Screening Challenge. *IEEE Trans. Med. Imaging* (2025). |
| SMDG | Kiefer, R. SMDG, A Standardized Fundus Glaucoma Dataset. Kaggle (2023). doi:10.34740/KAGGLE/DS/2329670 |
| AMC health-screen\* | Institutional cohort, Asan Medical Center — not externally published. |
| AMC clinic\* | Institutional cohort, Asan Medical Center — not externally published. |
**Downstream evaluation datasets**
| Dataset | Reference |
|---|---|
| APTOS 2019 | Karthik, Maggie & Dane, S. APTOS 2019 Blindness Detection. Kaggle (2019). |
| IDRiD | Porwal, P. et al. Indian Diabetic Retinopathy Image Dataset (IDRiD): A Database for Diabetic Retinopathy Screening Research. *Data* 3, 25 (2018). |
| MESSIDOR | Decencière, E. et al. Feedback on a Publicly Distributed Image Database: The MESSIDOR Database. *Image Anal. Stereol.* 231–234 (2014). |
| PAPILA | Kovalyk, O. et al. PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment. *Sci. Data* 9, 291 (2022). |
| Retina | jr2ngb. Retina Cataract Dataset. Kaggle (2016). |
| JSIEC | Cen, L.-P. et al. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. *Nat. Commun.* 12, 4828 (2021). |
| BRSET | Nakayama, L.F. et al. A Brazilian Multilabel Ophthalmological Dataset (BRSET). PhysioNet (2024). |
| ARIA | Farnell, D.J.J. et al. Enhancement of blood vessels in digital fundus photographs via the application of multiscale line operators. *J. Franklin Inst.* 345, 748–765 (2008). |
| MAPLES-DR | Lepetit-Aimon, G. et al. MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy. *Sci. Data* 11, 914 (2024). |
| AMC health-screen subset\* | Institutional cohort, Asan Medical Center — systemic/anthropometric & lifestyle biomarker fine-tuning. |
| AMC CAC cohort\* | Institutional cohort, Asan Medical Center — coronary artery calcium classification. |
## Citation
```bibtex
@article{lim_luna_2026,
title = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating
Clinical Metadata via Prediction of Demographics and Anatomical Structures},
author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo
and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
year = {2026},
}
```
## License
Released under the MIT License — Copyright (c) 2026 Hyunseok Lim.
The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center
institutional cohorts are not redistributable.