How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("black-forest-labs/FLUX.2-klein-base-4B", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("JeonghyeokDo/GeoSET")

prompt = "Turn this cat into a dog"
input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png")

image = pipe(image=input_image, prompt=prompt).images[0]

GeoSET

GeoSET is a generalist foundation model for SAR-to-EO image translation (SET). A pretrained text-to-image flow transformer is repurposed for SAR conditioning: the text stream is replaced by a spatial SAR stream fed by a speckle-robust SAR encoder, and the generator is pretrained on 3,204,744 curated SAR–EO pairs spanning diverse sensors, spatial resolutions and ground sampling distances. The resulting parent is adapted to downstream datasets by low-rank adaptation (LoRA), updating 0.60% of the generator parameters, or by full fine-tuning.

Files

Folder Content Size Licence
generalist/transformer GeoSETTransformer2DModel, Stage 2 generalist (500,000 updates, EMA), 3.852B parameters, float32 15.41 GB CC BY-NC-SA 4.0
generalist/sar_encoder GeoSETSAREncoder, Stage 1 speckle-robust SAR encoder (50,000 updates) 137.7 MB CC BY-NC-SA 4.0
generalist/vae AutoencoderFlux2, the frozen FLUX.2-klein-base-4B autoencoder converted to float32 336.2 MB Apache-2.0
lora/<dataset> one LoRA adapter per benchmark (rank 16, α 16, 23.10M parameters, EMA after 20,000 updates at batch size 16) 92.4 MB each CC BY-NC-SA 4.0
baselines/<dataset>/<method> 61 retrained comparison-method checkpoints – per method

generalist/ is a complete diffusers pipeline: model_index.json, the three model folders, scheduler/ and the pipeline code (pipeline.py). The same code files are also at the repository root and under src/geoset/ on GitHub. The full fine-tuned models are not released; the code that produces them is on GitHub.

Usage

import torch
from PIL import Image
from diffusers import DiffusionPipeline
from huggingface_hub import snapshot_download

root = snapshot_download("JeonghyeokDo/GeoSET", allow_patterns=["generalist/*", "lora/sar2opt/*"])
pipe = DiffusionPipeline.from_pretrained(f"{root}/generalist", custom_pipeline=f"{root}/generalist",
                                         torch_dtype=torch.float32).to("cuda")
pipe.load_lora("sar2opt")                                   # optional; resolved from {root}/lora/sar2opt
eo = pipe(Image.open("sar.png"), num_inference_steps=50, guidance_scale=2.0,
          generator=torch.Generator("cuda").manual_seed(20260812)).images[0]
eo.save("eo.png")

custom_pipeline points at the same folder because the pipeline, model and scheduler classes ship with the checkpoint. Keep torch_dtype=torch.float32; the pipeline casts the autoencoder to bfloat16 itself. The input is an 8-bit SAR image, read as one channel, or a list of equally sized images; every side must be a multiple of 16. Images are not resized or cropped inside the pipeline, so SAR2Opt tiles (600 × 600) are center-cropped to 512 × 512 first. The paper's results use 50 integration steps with two-pass classifier-free guidance of weight 2.0.

LoRA adapters

Adapter Benchmark Test pairs Input size
lora/qxs-saropt QXS-SAROPT 3,999 256 × 256
lora/sar2opt SAR2Opt 627 center 512 × 512 crop of the 600 × 600 tiles
lora/sar2eo SAR2EO first 4,000 of the 21,260 test pairs 256 × 256
lora/spacenet6 SpaceNet6 495 256 × 256

pipe.load_lora(name) merges one adapter into the transformer weights (name is one of the four keys above, or the path of an adapter directory). Without it, the pipeline runs the pretrained generalist. An adapter is merged once; load the pipeline again to switch to another one. The frozen train/test lists, the dataset preparation and scripts/translate.py, which translates a whole test split, are in the GitHub repository. Metrics of every checkpoint are in MODEL_ZOO.md.

Training

All stages train on 256 × 256 crops with AdamW and a gradient-norm clip of 1.0.

Stage 1 Stage 2 Stage 3, LoRA Stage 3, full FT
Trained SAR encoder all 3.852B generator parameters 23.10M adapter parameters (0.60%) all 3.852B generator parameters
Updates × global batch size 50,000 × 432 500,000 × 256 20,000 × 16 20,000 × 16
Learning rate 10⁻⁵, 500-update warmup 10⁻⁴ peak, 1,000-update warmup 10⁻⁴ 2 × 10⁻⁵
Hardware, time 4 × NVIDIA B200, 14.8 h 4 × NVIDIA B200, 129.3 h 1 GPU, 0.98 h per dataset 1 GPU, 1.34 h per dataset

Stage 1 trains the SAR encoder, initialised from the FLUX.2-klein-base-4B autoencoder encoder, to reconstruct SAR observations from speckle-perturbed copies through the frozen decoder. Stage 2 initialises the generator from FLUX.2-klein-base-4B, copies the image-stream parameters into the SAR stream, and learns conditional flow matching from Gaussian noise to the EO latent; the SAR condition is dropped with probability 0.1 so that two-pass classifier-free guidance can be used at inference. Stages 2 and 3 keep an EMA with decay 0.9999, and the released weights are the EMA parameters.

Training data. Pretraining: GUSO, TerraMesh, SARLO-80, SAR-1M and 3MOS, selected by the keep tables published on GitHub (corpus/keep_v1/). Adaptation: the training splits of QXS-SAROPT, SAR2Opt, SAR2EO and SpaceNet6. No dataset is redistributed here.

Comparison methods, in this same repository

baselines/ holds the sixteen prior methods of the paper's Tables 2 and 3, retrained on the same training splits and evaluated on the same test pairs: 61 checkpoints, one folder per dataset and method (Seg-CycleGAN on SpaceNet6 only), each with its own card and licence. Start at baselines/README.md; the training and inference code is in baselines/ on GitHub.

hf download JeonghyeokDo/GeoSET --include "baselines/sar2opt/c-diffset/*" --local-dir hf

Licences

This repository is mixed-licence, so the Hub tag is other.

  • GeoSET weights (generalist/transformer, generalist/sar_encoder, lora/): CC BY-NC-SA 4.0.
  • Autoencoder (generalist/vae): the FLUX.2-klein-base-4B autoencoder of Black Forest Labs, converted to float32, under Apache-2.0.
  • Code (the .py files here and the GitHub repository): Apache-2.0. The transformer and autoencoder code is adapted from black-forest-labs/flux2 (Apache-2.0).
  • Comparison methods (baselines/): the terms of the code each was trained with; the texts are in baselines/licenses/. Check the method you intend to use; the repository-level tag is not a substitute.

The generator is initialised from FLUX.2-klein-base-4B (Apache-2.0). GeoSET is not a FLUX product and is not endorsed by Black Forest Labs. See LICENSE-WEIGHTS.md and NOTICE on GitHub.

Citation

@article{do2026geoset,
  title={GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation},
  author={Do, Jeonghyeok and Kim, Munchurl},
  journal={arXiv preprint arXiv:2609.37496},
  year={2026}
}

Our prior work on SAR-to-EO image translation, C-DiffSET:

@article{do2026cdiffset,
  title={C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation},
  author={Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
  journal={IEEE Transactions on Circuits and Systems for Video Technology},
  year={2026},
  publisher={IEEE}
}
Downloads last month
-
Inference Providers NEW

Model tree for JeonghyeokDo/GeoSET

Adapter
(104)
this model

Paper for JeonghyeokDo/GeoSET