--- language: - en license: mit tags: - medical-imaging - ophthalmology - retinal-fundus - vision-language-model - multimodal - foundation-model - vl-bert - siglip2 - segmentation - zero-shot-classification pipeline_tag: image-feature-extraction library_name: pytorch --- # Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision **Luna** is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs). Instead of pixel reconstruction, it is pre-trained by **predicting masked clinical tokens** (disease label, age, sex, eye laterality) from the image, while a **UNETR decoder jointly supervises vessel / optic-disc segmentation**. Architecturally it is a single-stream VL-BERT: a **SigLIP-2 ViT-Base/16 (512px) vision encoder**, a **BERT-Base masked-LM initialized from PubMedBERT**, and a **UNETR head** for anatomy. > Paper: N/A > Code: https://github.com/loopback-kr/Luna ## Model Details | Attribute | Value | |---|---| | **Architecture** | SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder | | **Parameters** | 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse | | **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) | | **Vision encoder init.** | `google/siglip2-base-patch16-512` | | **Text encoder init.** | `NeuML/pubmedbert-base-embeddings` tokenizer, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens | | **Pre-training objective** | Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, *not* pixel-level MAE reconstruction | | **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) | | **License** | MIT | Prompt templates used during pre-training: `"A CFP image of {CLS}"`, `"age is {CLS} years old"`, `"gender is {CLS}"`, `"this is {CLS} direction eye"`. ## Files in this repository This is the full `accelerate` training state at epoch 399: | File | Contents | |---|---| | `model.safetensors` / `pytorch_model.bin` | Model weights (243,122,497 parameters, fp32) | | `optimizer.bin`, `scheduler.bin`, `scaler.pt` | AdamW / LR-scheduler / GradScaler state, for exact resumption | | `random_states_{0..3}.pkl` | RNG states of the four training processes | ## Usage Full training, fine-tuning and evaluation code lives in the [GitHub repository](https://github.com/loopback-kr/Luna) — see its `README.md` for the *Upstream training*, *Downstream training*, *Zero-Shot* and *Segmentation* sections. Checkpoint paths there expect `pytorch_model.bin`; passing the containing directory instead resumes optimizer and scheduler state as well. ## Training Data **\*** = private institutional data, not available externally. **Pre-training datasets** | Dataset | Reference | |---|---| | LAG | Li, L. et al. Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model. *CVPR* (2019). | | ODIR | Ocular Disease Intelligent Recognition (ODIR-2019). Peking University Grand Challenge (2019). | | PARAGUAY | Castillo Benítez, V.E. et al. Dataset from fundus images for the study of diabetic retinopathy. *Data in Brief* 36, 107068 (2021). | | G1020 | Bajwa, M.N. et al. G1020: A Benchmark Retinal Fundus Image Dataset for Computer-Aided Glaucoma Detection. *IJCNN* (2020). | | FUND-OCT | Hassan, T., Akram, M.U., Werghi, N. & Nazir, N. A hybrid convolutional framework for the automated extraction of retinal lesions and lesion-influenced grading of human retinal pathology. *IEEE J. Biomed. Health Inform.* 25(1) (2020). | | Drishti-GS1 | Sivaswamy, J. et al. Drishti-GS: Retinal Image Dataset for Optic Nerve Head Segmentation. *ISBI* 53–56 (2014). | | HRF | Budai, A. et al. Robust Vessel Segmentation in Fundus Images. *Int. J. Biomed. Imaging* 2013, 154860 (2013). | | ORIGA | Zhang, Z. et al. ORIGA(-light): An Online Retinal Fundus Image Database for Glaucoma Analysis and Research. *EMBC* 3065–3068 (2010). | | OIA-DDR | Li, T. et al. Diagnostic Assessment of Deep Learning Algorithms for Diabetic Retinopathy Screening. *Information Sciences* 501, 511–522 (2019). | | SUSTech-SYSU | Lin, L. et al. The SUSTech-SYSU dataset for automated exudate detection and diabetic retinopathy grading. *Sci. Data* 7, 409 (2020). | | JICHI | Takahashi, H. et al. Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy. *PLoS ONE* 12, e0179790 (2017). | | CHAKSU | Kumar J H, R. et al. Chákṣu: A glaucoma specific fundus image database. *Sci. Data* 10, 70 (2023). | | DR1 & DR2 | Pires, R. et al. Advancing Bag-of-Visual-Words Representations for Lesion Classification in Retinal Images. *PLoS ONE* 9, e96814 (2014). | | DeepDRiD | Liu, R. et al. DeepDRiD: Diabetic Retinopathy—Grading and Image Quality Estimation Challenge. *Patterns* 3, 100512 (2022). | | AIROGS-light-V2 | Kiefer, R. Glaucoma Dataset: EyePACS-AIROGS-light-V2. Kaggle (2024). doi:10.34740/KAGGLE/DSV/7802508 | | FIDVS | Jin, K. et al. Fundus Image Dataset for Vessel Segmentation. Kaggle (2025). doi:10.34740/KAGGLE/DS/7319687 | | JustRAIGS | Madadi, Y. et al. JustRAIGS: Justified Referral in AI Glaucoma Screening Challenge. *IEEE Trans. Med. Imaging* (2025). | | SMDG | Kiefer, R. SMDG, A Standardized Fundus Glaucoma Dataset. Kaggle (2023). doi:10.34740/KAGGLE/DS/2329670 | | AMC health-screen\* | Institutional cohort, Asan Medical Center — not externally published. | | AMC clinic\* | Institutional cohort, Asan Medical Center — not externally published. | **Downstream evaluation datasets** | Dataset | Reference | |---|---| | APTOS 2019 | Karthik, Maggie & Dane, S. APTOS 2019 Blindness Detection. Kaggle (2019). | | IDRiD | Porwal, P. et al. Indian Diabetic Retinopathy Image Dataset (IDRiD): A Database for Diabetic Retinopathy Screening Research. *Data* 3, 25 (2018). | | MESSIDOR | Decencière, E. et al. Feedback on a Publicly Distributed Image Database: The MESSIDOR Database. *Image Anal. Stereol.* 231–234 (2014). | | PAPILA | Kovalyk, O. et al. PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment. *Sci. Data* 9, 291 (2022). | | Retina | jr2ngb. Retina Cataract Dataset. Kaggle (2016). | | JSIEC | Cen, L.-P. et al. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. *Nat. Commun.* 12, 4828 (2021). | | BRSET | Nakayama, L.F. et al. A Brazilian Multilabel Ophthalmological Dataset (BRSET). PhysioNet (2024). | | ARIA | Farnell, D.J.J. et al. Enhancement of blood vessels in digital fundus photographs via the application of multiscale line operators. *J. Franklin Inst.* 345, 748–765 (2008). | | MAPLES-DR | Lepetit-Aimon, G. et al. MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy. *Sci. Data* 11, 914 (2024). | | AMC health-screen subset\* | Institutional cohort, Asan Medical Center — systemic/anthropometric & lifestyle biomarker fine-tuning. | | AMC CAC cohort\* | Institutional cohort, Asan Medical Center — coronary artery calcium classification. | ## Citation ```bibtex @article{lim_luna_2026, title = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating Clinical Metadata via Prediction of Demographics and Anatomical Structures}, author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug}, year = {2026}, } ``` ## License Released under the MIT License — Copyright (c) 2026 Hyunseok Lim. The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center institutional cohorts are not redistributable.