Update README.md
Browse files
README.md
CHANGED
|
@@ -11,30 +11,58 @@ tags:
|
|
| 11 |
- foundation-model
|
| 12 |
- vl-bert
|
| 13 |
- siglip2
|
| 14 |
-
|
|
|
|
|
|
|
| 15 |
library_name: pytorch
|
| 16 |
---
|
| 17 |
|
| 18 |
# Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
|
| 19 |
|
| 20 |
-
**Luna** is a multimodal, multi-task retinal foundation model
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
-
> Paper: *Multimodal Foundation Model of High-Resolution Fundus Photos with Clinical Metadata via Predicting Demographics and Anatomic Structure* — Lim, H. et al.
|
| 23 |
> Code: https://github.com/loopback-kr/Luna
|
| 24 |
|
| 25 |
## Model Details
|
| 26 |
|
| 27 |
| Attribute | Value |
|
| 28 |
|---|---|
|
| 29 |
-
| **Architecture** | SigLIP
|
| 30 |
-
| **Parameters** | 93.5M
|
| 31 |
| **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) |
|
| 32 |
-
| **Vision encoder init.** | `google/siglip2-base-patch16-512`
|
| 33 |
-
| **Text encoder init.** | `NeuML/pubmedbert-base-embeddings`, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens |
|
| 34 |
-
| **Pre-training objective** | Masked clinical-token prediction (label/age/sex/direction) + vessel
|
| 35 |
| **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) |
|
| 36 |
| **License** | MIT |
|
| 37 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
## Training Data
|
| 39 |
|
| 40 |
**\*** = private institutional data, not available externally.
|
|
@@ -84,8 +112,16 @@ library_name: pytorch
|
|
| 84 |
|
| 85 |
```bibtex
|
| 86 |
@article{lim_luna_2026,
|
| 87 |
-
title = {Multimodal Foundation Model
|
| 88 |
-
|
| 89 |
-
|
|
|
|
|
|
|
| 90 |
}
|
| 91 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
- foundation-model
|
| 12 |
- vl-bert
|
| 13 |
- siglip2
|
| 14 |
+
- segmentation
|
| 15 |
+
- zero-shot-classification
|
| 16 |
+
pipeline_tag: image-feature-extraction
|
| 17 |
library_name: pytorch
|
| 18 |
---
|
| 19 |
|
| 20 |
# Luna — Retinal Foundation Model with Hybrid Clinical & Anatomical Supervision
|
| 21 |
|
| 22 |
+
**Luna** is a multimodal, multi-task retinal foundation model for high-resolution color fundus photographs (CFPs).
|
| 23 |
+
Instead of pixel reconstruction, it is pre-trained by **predicting masked clinical tokens** (disease label, age, sex,
|
| 24 |
+
eye laterality) from the image, while a **UNETR decoder jointly supervises vessel / optic-disc segmentation**.
|
| 25 |
+
|
| 26 |
+
Architecturally it is a single-stream VL-BERT: a **SigLIP-2 ViT-Base/16 (512px) vision encoder**, a **BERT-Base
|
| 27 |
+
masked-LM initialized from PubMedBERT**, and a **UNETR head** for anatomy.
|
| 28 |
+
|
| 29 |
+
> Paper: N/A
|
| 30 |
|
|
|
|
| 31 |
> Code: https://github.com/loopback-kr/Luna
|
| 32 |
|
| 33 |
## Model Details
|
| 34 |
|
| 35 |
| Attribute | Value |
|
| 36 |
|---|---|
|
| 37 |
+
| **Architecture** | SigLIP-2 ViT-Base/16 vision encoder + BERT-Base masked-LM (single-stream VL-BERT fusion, 128 compressed vision tokens) + UNETR segmentation decoder |
|
| 38 |
+
| **Parameters** | 243M total; ≈93.5M in the SigLIP-2 ViT-Base vision encoder that downstream tasks reuse |
|
| 39 |
| **Input resolution** | 512 × 512 RGB (CLAHE-enhanced) |
|
| 40 |
+
| **Vision encoder init.** | `google/siglip2-base-patch16-512` |
|
| 41 |
+
| **Text encoder init.** | `NeuML/pubmedbert-base-embeddings` tokenizer, extended with `[LABEL]`, `[AGE]`, `[SEX]`, `[DIR]` special tokens |
|
| 42 |
+
| **Pre-training objective** | Masked clinical-token prediction (label / age / sex / direction) + vessel & optic-disc Dice loss, *not* pixel-level MAE reconstruction |
|
| 43 |
| **Pre-training data** | 361,519 CFPs (52K public + 309K private institutional images) |
|
| 44 |
| **License** | MIT |
|
| 45 |
|
| 46 |
+
Prompt templates used during pre-training: `"A CFP image of {CLS}"`, `"age is {CLS} years old"`,
|
| 47 |
+
`"gender is {CLS}"`, `"this is {CLS} direction eye"`.
|
| 48 |
+
|
| 49 |
+
## Files in this repository
|
| 50 |
+
|
| 51 |
+
This is the full `accelerate` training state at epoch 399:
|
| 52 |
+
|
| 53 |
+
| File | Contents |
|
| 54 |
+
|---|---|
|
| 55 |
+
| `model.safetensors` / `pytorch_model.bin` | Model weights (243,122,497 parameters, fp32) |
|
| 56 |
+
| `optimizer.bin`, `scheduler.bin`, `scaler.pt` | AdamW / LR-scheduler / GradScaler state, for exact resumption |
|
| 57 |
+
| `random_states_{0..3}.pkl` | RNG states of the four training processes |
|
| 58 |
+
|
| 59 |
+
## Usage
|
| 60 |
+
|
| 61 |
+
Full training, fine-tuning and evaluation code lives in the
|
| 62 |
+
[GitHub repository](https://github.com/loopback-kr/Luna) — see its `README.md` for the *Upstream training*,
|
| 63 |
+
*Downstream training*, *Zero-Shot* and *Segmentation* sections. Checkpoint paths there expect
|
| 64 |
+
`pytorch_model.bin`; passing the containing directory instead resumes optimizer and scheduler state as well.
|
| 65 |
+
|
| 66 |
## Training Data
|
| 67 |
|
| 68 |
**\*** = private institutional data, not available externally.
|
|
|
|
| 112 |
|
| 113 |
```bibtex
|
| 114 |
@article{lim_luna_2026,
|
| 115 |
+
title = {Multimodal Foundation Model for High-Resolution Fundus Photographs Incorporating
|
| 116 |
+
Clinical Metadata via Prediction of Demographics and Anatomical Structures},
|
| 117 |
+
author = {Lim, Hyunseok and Kim, Junseok and Oh, Joonseo and Kim, Kanghyun and Lim, Jongsoo
|
| 118 |
+
and Jeong, Jinhoon and Jeong, Hy and Kim, Yoonjeon and Kim, Namkug},
|
| 119 |
+
year = {2026},
|
| 120 |
}
|
| 121 |
```
|
| 122 |
+
|
| 123 |
+
## License
|
| 124 |
+
|
| 125 |
+
Released under the MIT License — Copyright (c) 2026 Hyunseok Lim.
|
| 126 |
+
The third-party datasets listed above carry their own licenses and access terms, and the Asan Medical Center
|
| 127 |
+
institutional cohorts are not redistributable.
|