FLUX.2 Klein Alpha: an RGBA VAE with extraction and removal LoRAs
Xavier Jara · tr-mz.com
Image generators like FLUX.2 Klein produce opaque images because their VAE has three colour channels. This repo contains a FLUX.2 VAE widened to four channels (RGBA) and fine-tuned to encode and decode transparency. It also contains two LoRAs for FLUX.2 Klein Base 9B that use it:
| File | What it is | Size |
|---|---|---|
vae/ |
RGBA VAE (diffusers AutoencoderKLFlux2, 4 in / 4 out channels, 32 latent channels), 58,000 fine-tuning steps |
336 MB |
loras/extract_9b.safetensors |
Extract-9B: pulls an object out of a picture as a transparent RGBA image | 166 MB |
loras/remove_9b.safetensors |
Remove-9B: erases an object and fills in the background | 331 MB |
Licence: non-commercial only. See License and notices.
Example
Extract-9B gets the photo and the same wall without the hands, and returns the hands and their shadow as one RGBA layer. The shadow comes out as a soft, semi-transparent dark layer rather than a solid shape. The transparent result is assets/example_hand_extract_9b.png.
Photo: Shutterstock #643993384 (via KiwiCo). Background plate generated with Google Gemini from the same photo.
How the models are used
Extract-9B takes two images:
- the composite: the object on a background, which is the picture you want to cut the object out of;
- the background plate: the same background without the object.
It returns the object as an RGBA image at the same size and position. Soft details such as shadows, glows and semi-transparent parts come out as partial alpha.
Remove-9B takes one image: the photo as RGBA, with alpha 128 on the object to remove and 255 everywhere else. It returns the photo with the object removed, usually including its shadow.
The two chain together. Remove-9B can make the background plate that Extract-9B needs, so a user only has to brush over the object.
Recommended settings:
| Prompt | Steps | Guidance | |
|---|---|---|---|
| Extract-9B | Foreground |
6 | 4 |
| Remove-9B | Photo with the object removed, clean background preserved. |
30 | 4 |
Both LoRAs were trained with a modified ai-toolkit. They need a sampler that feeds four-channel control images and decodes with the RGBA VAE. The code for that, a patch to ai-toolkit and a Gradio demo are in the project repository (being prepared for release).
Extract-9B works at 6 steps. On the 20-image test set it scores the same at 6 steps as at 25 (IoU 0.934 vs 0.934; alpha error 0.0096 vs 0.0090), about 4× faster. Remove-9B has only been checked at 30 steps.
On one 16 GB GPU, with the 9B base in fp8, expect about 10 GB of VRAM. Extract-9B takes about 11 s per image at 6 steps (about 46 s at 25), and Remove-9B about 40 s at 30 steps, at 512².
Download
huggingface-cli download trmz/flux2-klein-alpha --local-dir flux2-klein-alpha
The base model is black-forest-labs/FLUX.2-klein-base-9B. Accept its licence on Hugging Face first.
Loading only the VAE:
from diffusers import AutoencoderKLFlux2
vae = AutoencoderKLFlux2.from_pretrained("trmz/flux2-klein-alpha", subfolder="vae")
# input: RGBA tensor in [-1, 1], shape (B, 4, H, W); output: RGBA in the same range
Results
VAE, on 24 held-out RGBA images at 256 px. Values are RMSE on a 0–1 scale, so lower is better.
| Alpha | Colour (flattened on black) | Colour (flattened on white) | |
|---|---|---|---|
| Stock FLUX.2 VAE (RGB only) | – | 0.0121 | 0.0138 |
| This VAE | 0.0204 | 0.0133 | 0.0163 |
Adding transparency costs a little colour accuracy compared with the original VAE.
Extract-9B, on 20 held-out transparent objects, counting a pixel as object when alpha > 50%:
| Value | |
|---|---|
| IoU | 0.934 |
| Edge IoU (5 px band) | 0.827 |
| Mean alpha error | 0.009 |
On the same test, a full fine-tune of Klein 4B reached IoU 0.879.
Limitations
- Remove-9B works well on photos. On flat graphics and illustrations it can fill the hole with odd colours.
- Small tests: 24 images for the VAE, 20 for extraction. Treat small differences as noise.
- Extract-9B needs a background plate. Without one, generate it with Remove-9B first.
- Resolutions above 512² are untested.
Training data
- VAE: 5,829 real transparent (RGBA) images, with the triple-background loss from AlphaVAE (the image, plus it flattened onto black and onto white).
- Extract-9B: 2,039 synthetic composites of transparent objects placed on backgrounds, with the background plates.
- Remove-9B: the OBER dataset from ObjectClear (S-Lab License 1.0, non-commercial).
License and notices
- All three weights fall under the FLUX Non-Commercial License: the VAE is derived from the FLUX.2 VAE (FLUX.2-dev), and the LoRAs are trained for FLUX.2 Klein Base 9B.
- Remove-9B is also subject to the S-Lab License 1.0 of its training data.
- No commercial use is permitted.
- The full attribution notices are in NOTICE.md.
This is an independent research project, not endorsed by Black Forest Labs, S-Lab or Qwen.
Credits
- AlphaVAE: Wang, Yu, Zhan, Yuan, 2025, arXiv:2507.09308. The RGBA VAE method and composite loss.
- ObjectClear / OBER dataset: Zhao, Zhou, Loy et al., arXiv:2505.22636.
- ai-toolkit by ostris (MIT): the trainer, which I modified.
- LayerDiffuse and its FLUX.1 adaptation: the idea of pairing a transparent VAE with adapters.
- FLUX.2 by Black Forest Labs.
Model tree for trmz/flux2-klein-alpha
Base model
black-forest-labs/FLUX.2-klein-base-9B