Robotics
Diffusers
goal-image-generation
image-editing
lora
flux

SGE-Goal — LoRA checkpoints

LoRA adapters for SGE-Goal, from Goal Images Are Control Targets: Spatially-Grounded Goal-Image Synthesis for Robotic Manipulation (WACV 2027).

SGE-Goal synthesizes a goal image for a goal-conditioned manipulation policy in three steps: a training-free initial edit-region mask M_i (Step 1), a fine-tuned final edit-region mask M_f (Step 2) and the fine-tuned goal image Î_e (Step 3). This repository holds the Step 2 and Step 3 adapters for both base editors and all three datasets, plus the mask-ablation adapters.

Code, data pipeline, evaluator and usage instructions: https://github.com/Riddhi-Chatterjee/SGE-Goal

Contents

Path Base editor Dataset Stage Epoch
flux1-kontext-dev/bridge/step2_final_mask/ FLUX.1 Kontext [dev] BridgeDataV2 Step 2 · final mask 38
flux1-kontext-dev/calvin/step2_final_mask/ FLUX.1 Kontext [dev] CALVIN-ABCD Step 2 · final mask 22
flux1-kontext-dev/libero/step2_final_mask/ FLUX.1 Kontext [dev] LIBERO-10/90 Step 2 · final mask 56
flux1-kontext-dev/bridge/step3_goal_image/ FLUX.1 Kontext [dev] BridgeDataV2 Step 3 · goal image 18
flux1-kontext-dev/calvin/step3_goal_image/ FLUX.1 Kontext [dev] CALVIN-ABCD Step 3 · goal image 18
flux1-kontext-dev/libero/step3_goal_image/ FLUX.1 Kontext [dev] LIBERO-10/90 Step 3 · goal image 64
flux2-klein/bridge/step2_final_mask/ FLUX.2 klein BridgeDataV2 Step 2 · final mask 17
flux2-klein/calvin/step2_final_mask/ FLUX.2 klein CALVIN-ABCD Step 2 · final mask 22
flux2-klein/libero/step2_final_mask/ FLUX.2 klein LIBERO-10/90 Step 2 · final mask 59
flux2-klein/bridge/step3_goal_image/ FLUX.2 klein BridgeDataV2 Step 3 · goal image 9
flux2-klein/calvin/step3_goal_image/ FLUX.2 klein CALVIN-ABCD Step 3 · goal image 15
flux2-klein/libero/step3_goal_image/ FLUX.2 klein LIBERO-10/90 Step 3 · goal image 71
ablations/flux1-kontext-dev/bridge/wo_final_mask/ FLUX.1 Kontext [dev] BridgeDataV2 Step 3, without M_f 9
ablations/flux1-kontext-dev/bridge/wo_masks/ FLUX.1 Kontext [dev] BridgeDataV2 Step 3, without M_i, M_f 13
ablations/flux2-klein/bridge/wo_final_mask/ FLUX.2 klein BridgeDataV2 Step 3, without M_f 9
ablations/flux2-klein/bridge/wo_masks/ FLUX.2 klein BridgeDataV2 Step 3, without M_i, M_f 10

Each directory holds the adapter weights (pytorch_lora_weights.safetensors, or transformer_lora.safetensors for the FLUX.2 klein Step 2 adapters) and, where available, the trainer's state file recording the epoch and step. Optimizer and scheduler states are not included. SHA256SUMS lists the checksum of every weight file.

Training: LoRA rank 16, α 16, 8-bit AdamW, learning rate 1e-4 (constant, 200 warm-up steps), batch size 160, bf16; 256 px for BridgeDataV2 and CALVIN, 128 px for LIBERO. Step 2 adds a channel-consistency loss (λ = 25) and Step 3 a static-scene-consistency loss (λ = 7), both applied for σ < 0.35.

Usage

huggingface-cli download RiddhiCh/SGE-Goal --local-dir sge-goal-checkpoints
sha256sum -c sge-goal-checkpoints/SHA256SUMS --ignore-missing   # run from inside the folder

Then point the pipeline drivers in the code repository at the matching directories, e.g. for FLUX.2 klein on BridgeDataV2:

cd sge_goal/pipelines/flux2_klein
python run_pipeline.py --data-dir /abs/path/to/test_data \
  --fer-checkpoint-dir /abs/path/sge-goal-checkpoints/flux2-klein/bridge/step2_final_mask \
  --ie-checkpoint-dir  /abs/path/sge-goal-checkpoints/flux2-klein/bridge/step3_goal_image

Licence

The adapters are released by the authors for research use. They are fine-tuned from FLUX.1 Kontext [dev] and FLUX.2 klein; using them requires the base model and therefore complying with each base model's licence (FLUX.1 [dev] is under the FLUX.1 [dev] Non-Commercial License). The licences of the training datasets (BridgeData V2, CALVIN, LIBERO) also apply.

Citation

The paper has been accepted at WACV 2027; this entry will be updated once the proceedings are published.

@inproceedings{chatterjee2027goalimages,
  title     = {Goal Images Are Control Targets: Spatially-Grounded Goal-Image Synthesis for Robotic Manipulation},
  author    = {Chatterjee, Riddhi and Shrirao, Ritish and Mehrotra, Kushagra and Gopalakrishnan, Viswanath},
  booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
  year      = {2027}
}
Downloads last month
-
Video Preview
loading

Model tree for RiddhiCh/SGE-Goal

Adapter
(246)
this model

Datasets used to train RiddhiCh/SGE-Goal