SimpleTuner Training Assistant v1 for Qwen Image 2.1
Built with Qwen. This experimental training assistant LoRA was trained from scratch for 1,000 updates, using corrected adamw_bf16, learning rate 1e-4 and global gradient norm clipping at 1.0. It replaces the earlier v1 weights in this repository.
Keep this adapter frozen and active during downstream LoRA training, then disable it when generating images with the downstream adapter. Read the evaluation below before choosing it for an experiment.
Use in SimpleTuner
Use a SimpleTuner checkout with Qwen Image 2.1 and assistant-LoRA support. Add these fields to a complete concept-training configuration:
{
"model_family": "qwen_image",
"model_flavour": "v2.1",
"model_type": "lora",
"lora_type": "standard",
"assistant_lora_path": "SimpleTuner/Qwen-Image-2.1-training-assistant-v1",
"assistant_lora_weight_name": "pytorch_lora_weights.safetensors",
"disable_assistant_lora": false,
"assistant_lora_strength": 1.0,
"assistant_lora_inference_strength": 0.0
}
SimpleTuner freezes the assistant alongside the trainable concept adapter. Use ordinary concept training; distillation_method: assistant_lora creates an assistant using online teacher generation. At inference, load the base model and your resulting concept LoRA with this assistant disabled. This assistant has no trigger word. It is not a character adapter or a few-step generation adapter. Older Qwen Image flavours have not been tested with these weights.
Training
Training used the reusable set of 10,000 generated images, pinned to dataset revision 7fb60061955890fbcc697214cbfe9f9f5fd44ce9. Eight H100 GPUs generated the images from the unmodified Qwen Image 2.1 base model with 40 native inference steps, CFG 1 and BF16. Captions use CC12M's long_caption field; original CC12M images were not used as targets.
The images were decoded with the full-frame VAE and stored as PNGs. Training re-encodes those images, so the targets differ from directly cached terminal teacher latents. The L40S run used extracted copies of the published shards. The SimpleTuner example qwen_image-2.1-assistant-lora-offline.peft-lora reads the same data through Webshart.
| Setting | Value |
|---|---|
| GPU | NVIDIA L40S |
| Updates / batch / accumulation | 1,000 / 1 / 1 |
| Precision | BF16 |
| LoRA rank / alpha | 32 / 32 |
| Attention projections | to_q, to_k, to_v, to_out.0 |
| Optimizer | Corrected adamw_bf16 |
| Learning rate | 1e-4, constant after 25 warmup steps |
| Gradient clipping | Global norm, maximum 1.0 |
| Gradient checkpointing | Enabled, interval 2 |
| Sampling | Equal weight for each of 12 image backends; no repeats |
| VAE | Full-frame encode/decode; encoding batch 1 |
| Trainable parameters | 33,554,432 |
Training samples square, portrait and landscape buckets at four base resolutions:
| Base resolution | Square | Portrait | Landscape |
|---|---|---|---|
| 512 | 512 × 512 | 384 × 672 | 672 × 384 |
| 1024 | 1024 × 1024 | 768 × 1344 | 1344 × 768 |
| 1536 | 1536 × 1536 | 1152 × 2016 | 2016 × 1152 |
| 2048 | 2048 × 2048 | 1536 × 2688 | 2688 × 1536 |
The run budgets 1,000 sampled images from the 10,000-image pool; it does not make a complete pass over that dataset. The nominal 1024x1024 field in the weights metadata is not the complete list of training sizes. See training_details.json for provenance.
Evaluation and limitations
This is an experimental assistant; successful Domokun concept training was not demonstrated.
The assistant itself was visually checked on four prompts at 1024 and 2048 pixels after 1,000 updates. The requested subjects remained recognizable, with composition and style changes; bicycle geometry artifacts remained. These checks do not establish broad quality preservation.
A downstream comparison trained fresh rank-32 concept adapters on RareConcepts/Domokun for 1,000 updates at 2048 pixels, using the brown-square emoji trigger (🟫). Both runs used corrected adamw_bf16, learning rate 1e-4, global norm clipping at 1.0, batch size 1, matching initial concept weights and initial training RNG states. Baseline validation images were pixel-identical. The control continued from an audited step-250 checkpoint to step 1,000 with unchanged data and batch settings. The assisted run used this frozen assistant at strength 1.0 during training and disabled it for every validation. Validation used 40 native inference steps and CFG 1.
At update 1,000:
- Neither run reliably produced Domokun as the subject of the beach or red-scarf prompts; the main subjects were still people.
- The control introduced strong Domokun-like features into an unrelated fox prompt: a brown rectangular creature with a red mouth and triangular teeth. The assisted run retained a recognizable orange fox reading a book, with a substantial shift toward cartoon illustration.
- An unrelated elderly-woman portrait remained coherent in both runs.
This small comparison suggests reduced concept spillover on the tested fox prompt, but does not demonstrate a successful concept-learning recipe or a general downstream quality benefit. It covers one training seed, two character prompts and two unrelated prompts. Do not interpret preserved unrelated subjects alone as evidence that the new concept has been learned.
The optimizer's stochastic addition now computes input + alpha * other. The earlier v1 used the reversed coefficient placement, a learning rate of 1e-5, and individual-value clipping at 0.01. This release starts with fresh adapter and optimizer state. These changes do not by themselves establish a downstream quality benefit.
Artifact and license
pytorch_lora_weights.safetensors contains 256 finite BF16 tensors and matches the completed checkpoint at update 1,000.
SHA-256: 2571f7fa98824bee87b7fc2a8d6adf7347a9e9032e5b5fc5908e5ac76ce87f9d.
Modification notice: this separately trained LoRA changes Qwen Image 2.1 transformer attention projections when loaded. Distribution and use are subject to the Qwen Research License Agreement, including its non-commercial terms. See Notice for upstream attribution.
Model tree for SimpleTuner/Qwen-Image-2.1-training-assistant-v1
Base model
Qwen/Qwen-Image-2.1