Instructions to use JeonghyeokDo/GeoSET with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use JeonghyeokDo/GeoSET with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("JeonghyeokDo/GeoSET") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("JeonghyeokDo/GeoSET")
prompt = "Turn this cat into a dog"
input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png")
image = pipe(image=input_image, prompt=prompt).images[0]GeoSET
GeoSET is a generalist foundation model for SAR-to-EO image translation (SET). A pretrained text-to-image flow transformer is repurposed for SAR conditioning: the text stream is replaced by a spatial SAR stream fed by a speckle-robust SAR encoder, and the generator is pretrained on 3,204,744 curated SAR–EO pairs spanning diverse sensors, spatial resolutions and ground sampling distances. The resulting parent is adapted to downstream datasets by low-rank adaptation (LoRA), updating 0.60% of the generator parameters, or by full fine-tuning.
- Paper: https://arxiv.org/abs/2609.37496
- Code: https://github.com/KAIST-VICLab/GeoSET
- Project page: https://kaist-viclab.github.io/GeoSET_site/
- Comparison-method weights:
baselines/in this repository: the sixteen prior methods of the paper's Tables 2 and 3 on the four public benchmarks
Files
| Folder | Content | Size | Licence |
|---|---|---|---|
generalist/transformer |
GeoSETTransformer2DModel, Stage 2 generalist (500,000 updates, EMA), 3.852B parameters, float32 |
15.41 GB | CC BY-NC-SA 4.0 |
generalist/sar_encoder |
GeoSETSAREncoder, Stage 1 speckle-robust SAR encoder (50,000 updates) |
137.7 MB | CC BY-NC-SA 4.0 |
generalist/vae |
AutoencoderFlux2, the frozen FLUX.2-klein-base-4B autoencoder converted to float32 |
336.2 MB | Apache-2.0 |
lora/<dataset> |
one LoRA adapter per benchmark (rank 16, α 16, 23.10M parameters, EMA after 20,000 updates at batch size 16) | 92.4 MB each | CC BY-NC-SA 4.0 |
baselines/<dataset>/<method> |
61 retrained comparison-method checkpoints | – | per method |
generalist/ is a complete diffusers pipeline: model_index.json, the three model folders, scheduler/
and the pipeline code (pipeline.py). The same code files are also at the repository root and under
src/geoset/ on GitHub. The full fine-tuned models are not released; the code that produces them is on GitHub.
Usage
import torch
from PIL import Image
from diffusers import DiffusionPipeline
from huggingface_hub import snapshot_download
root = snapshot_download("JeonghyeokDo/GeoSET", allow_patterns=["generalist/*", "lora/sar2opt/*"])
pipe = DiffusionPipeline.from_pretrained(f"{root}/generalist", custom_pipeline=f"{root}/generalist",
torch_dtype=torch.float32).to("cuda")
pipe.load_lora("sar2opt") # optional; resolved from {root}/lora/sar2opt
eo = pipe(Image.open("sar.png"), num_inference_steps=50, guidance_scale=2.0,
generator=torch.Generator("cuda").manual_seed(20260812)).images[0]
eo.save("eo.png")
custom_pipeline points at the same folder because the pipeline, model and scheduler classes ship with the
checkpoint. Keep torch_dtype=torch.float32; the pipeline casts the autoencoder to bfloat16 itself. The
input is an 8-bit SAR image, read as one channel, or a list of equally sized images; every side must be a
multiple of 16. Images are not resized or cropped inside the pipeline, so SAR2Opt tiles (600 × 600) are
center-cropped to 512 × 512 first. The paper's results use 50 integration steps with two-pass classifier-free guidance of
weight 2.0.
LoRA adapters
| Adapter | Benchmark | Test pairs | Input size |
|---|---|---|---|
lora/qxs-saropt |
QXS-SAROPT | 3,999 | 256 × 256 |
lora/sar2opt |
SAR2Opt | 627 | center 512 × 512 crop of the 600 × 600 tiles |
lora/sar2eo |
SAR2EO | first 4,000 of the 21,260 test pairs | 256 × 256 |
lora/spacenet6 |
SpaceNet6 | 495 | 256 × 256 |
pipe.load_lora(name) merges one adapter into the transformer weights (name is one of the four keys above,
or the path of an adapter directory). Without it, the pipeline runs the pretrained generalist. An adapter is
merged once; load the pipeline again to switch to another one. The frozen train/test lists, the dataset
preparation and scripts/translate.py, which translates a whole test split, are in the
GitHub repository. Metrics of every checkpoint are in
MODEL_ZOO.md.
Training
All stages train on 256 × 256 crops with AdamW and a gradient-norm clip of 1.0.
| Stage 1 | Stage 2 | Stage 3, LoRA | Stage 3, full FT | |
|---|---|---|---|---|
| Trained | SAR encoder | all 3.852B generator parameters | 23.10M adapter parameters (0.60%) | all 3.852B generator parameters |
| Updates × global batch size | 50,000 × 432 | 500,000 × 256 | 20,000 × 16 | 20,000 × 16 |
| Learning rate | 10⁻⁵, 500-update warmup | 10⁻⁴ peak, 1,000-update warmup | 10⁻⁴ | 2 × 10⁻⁵ |
| Hardware, time | 4 × NVIDIA B200, 14.8 h | 4 × NVIDIA B200, 129.3 h | 1 GPU, 0.98 h per dataset | 1 GPU, 1.34 h per dataset |
Stage 1 trains the SAR encoder, initialised from the FLUX.2-klein-base-4B autoencoder encoder, to reconstruct SAR observations from speckle-perturbed copies through the frozen decoder. Stage 2 initialises the generator from FLUX.2-klein-base-4B, copies the image-stream parameters into the SAR stream, and learns conditional flow matching from Gaussian noise to the EO latent; the SAR condition is dropped with probability 0.1 so that two-pass classifier-free guidance can be used at inference. Stages 2 and 3 keep an EMA with decay 0.9999, and the released weights are the EMA parameters.
Training data. Pretraining: GUSO, TerraMesh, SARLO-80, SAR-1M and 3MOS, selected by the keep tables
published on GitHub (corpus/keep_v1/). Adaptation: the training splits of QXS-SAROPT, SAR2Opt, SAR2EO and
SpaceNet6. No dataset is redistributed here.
Comparison methods, in this same repository
baselines/ holds the sixteen prior methods of the paper's Tables 2 and 3, retrained on the same training
splits and evaluated on the same test pairs: 61 checkpoints, one folder per dataset and method (Seg-CycleGAN on
SpaceNet6 only), each with its own card and licence. Start at
baselines/README.md; the
training and inference code is in
baselines/ on GitHub.
hf download JeonghyeokDo/GeoSET --include "baselines/sar2opt/c-diffset/*" --local-dir hf
Licences
This repository is mixed-licence, so the Hub tag is other.
- GeoSET weights (
generalist/transformer,generalist/sar_encoder,lora/): CC BY-NC-SA 4.0. - Autoencoder (
generalist/vae): the FLUX.2-klein-base-4B autoencoder of Black Forest Labs, converted to float32, under Apache-2.0. - Code (the
.pyfiles here and the GitHub repository): Apache-2.0. The transformer and autoencoder code is adapted from black-forest-labs/flux2 (Apache-2.0). - Comparison methods (
baselines/): the terms of the code each was trained with; the texts are inbaselines/licenses/. Check the method you intend to use; the repository-level tag is not a substitute.
The generator is initialised from FLUX.2-klein-base-4B (Apache-2.0). GeoSET is not a FLUX product and is not endorsed by Black Forest Labs. See LICENSE-WEIGHTS.md and NOTICE on GitHub.
Citation
@article{do2026geoset,
title={GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation},
author={Do, Jeonghyeok and Kim, Munchurl},
journal={arXiv preprint arXiv:2609.37496},
year={2026}
}
Our prior work on SAR-to-EO image translation, C-DiffSET:
@article{do2026cdiffset,
title={C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation},
author={Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
journal={IEEE Transactions on Circuits and Systems for Video Technology},
year={2026},
publisher={IEEE}
}
- Downloads last month
- -
Model tree for JeonghyeokDo/GeoSET
Base model
black-forest-labs/FLUX.2-klein-base-4B