SimpleTuner Training Assistant v3 for Qwen Image 2.1
Built with Qwen. This training assistant LoRA was trained from scratch for 30,000 updates, with rank/alpha 64, four base resolutions, and flow schedule auto shift. It uses corrected adamw_bf16, learning rate 1e-4, and global gradient norm clipping at 1.0.
Keep this adapter frozen and active during downstream LoRA training, then disable it when generating images with the downstream adapter. Its purpose is to help preserve coherence and image quality during fine-tuning.
Compared with v2, v3 increases rank from 32 to 64 and assistant training from 1,000 to 30,000 updates, uses a larger data pool and effective batch size, and enables automatic flow-schedule shifting. V3 starts from a fresh adapter; it does not resume or expand v2. V1 used synthetic targets only; v2 and v3 also use real images.
Use in SimpleTuner
Use a SimpleTuner checkout with Qwen Image 2.1 and assistant-LoRA support. Add these fields to a complete concept-training configuration:
{
"model_family": "qwen_image",
"model_flavour": "v2.1",
"model_type": "lora",
"lora_type": "standard",
"assistant_lora_path": "SimpleTuner/Qwen-Image-2.1-training-assistant-v3",
"disable_assistant_lora": false,
"assistant_lora_strength": 1.0,
"assistant_lora_inference_strength": 0.0
}
SimpleTuner freezes the assistant alongside the trainable concept adapter. Use ordinary concept training; distillation_method: assistant_lora creates an assistant using online teacher generation. At inference, load the base model and your resulting concept LoRA with this assistant disabled. This assistant has no trigger word. It is not a character adapter or a few-step generation adapter. Older Qwen Image flavours have not been tested with these weights.
The assistant is intentionally trained to absorb changes that would otherwise affect downstream tuning. Images made with the assistant enabled by itself do not represent the intended final inference setup. The relevant evaluation is downstream training with the assistant enabled and frozen, followed by inference with it disabled.
Training
The source pool contains 30,000 unique images, divided equally among:
- The full 10,000 model-generated images, generated from Qwen Image 2.1 with 40 inference steps, CFG 1 and BF16, using CC12M captions.
- 10,000 real CC12M images, accessed through structured-caption indices, using
long_caption. - 10,000 real e621 images, accessed through Webshart indices.
Each source group has one third of the sampling probability. Synthetic images retain their generated dimensions in 12 aspect/resolution backends. The same real-image IDs are reused across four resolution backends per source; these are not additional unique images. All eight ranks' completed caches were audited against these counts.
| Setting | Value |
|---|---|
| Hardware | 8 ร NVIDIA H100 |
| Updates | 30,000 |
| Batch per process / accumulation / effective batch | 2 / 1 / 16 |
| Precision | BF16 |
| LoRA rank / alpha | 64 / 64 |
| Attention projections | to_q, to_k, to_v, to_out.0 |
| Trainable parameters | 67,108,864 |
| Optimizer | Corrected adamw_bf16 |
| Learning rate | 1e-4, constant after 25 warmup steps |
| Gradient clipping | Global norm, maximum 1.0 |
| Gradient checkpointing | Enabled, interval 1 |
| Base resolutions | 512, 1024, 1536, 2048; aspect buckets |
| Flow schedule | Auto shift enabled |
| REPA / regularisation | Neither used |
| VAE | Original Qwen Image 2.1; tiling disabled, encoding batch 1 |
The resumed phase after checkpoint 1,000 recorded 29,000 sampled batches: 9,587 synthetic, 9,676 CC12M and 9,737 e621. Data are sampled repeatedly over 30,000 updates. The nominal 512x512 value in the weights metadata is not the complete list of training sizes.
See training_details.json for artifact provenance and resume notes, training_config.json and dataloader.json for the recipe, and prompts.json for validation prompts. The published config resets the resume setting for a fresh run.
Final assistant validation
The comparisons below show the bare base model and base + v3 assistant at update 30,000, with assistant strength 1.0, 40 inference steps, CFG 1 and validation seed 42. Columns are output resolutions. Original comparison grids were split into separate, full-resolution panels without resizing; the bottom prompt strip and dividing gutter were removed, while the embedded model labels were retained.
These show what the assistant itself changes. They do not show a downstream concept LoRA trained with v3 and then evaluated with the assistant disabled. V3's downstream benefit and its comparison with v2 remain to be evaluated. Existing LoRA experiments establish the context for earlier assistants, not a v3 result.
Samples use the original Qwen Image 2.1 VAE. Texture differences from Ollin's texture-fixed decoder are separate from the assistant's effect.
Expand: portrait โ base versus assistant v3 at 30k
A photograph of an elderly woman wearing a blue silk scarf, soft window light, a plain gray background.
Expand: street โ base versus assistant v3 at 30k
A photograph of a woman walking down a rain-soaked city street, reflections in the pavement, natural evening light.
Expand: interior โ base versus assistant v3 at 30k
A photograph of a sunlit kitchen with wooden cabinets, ceramic cups and fresh flowers on a table.
Expand: landscape โ base versus assistant v3 at 30k
A photograph of a mountain lake at dawn with mist over the water and pine trees along the shore.
Expand: fabric โ base versus assistant v3 at 30k
A close-up photograph of a white cotton shirt and a folded blue silk scarf on a wooden table.
Expand: fox โ base versus assistant v3 at 30k
An illustration of an anthropomorphic fox wearing a blue jacket and reading a book in a cozy library.
Artifact and license
pytorch_lora_weights.safetensors contains 256 finite BF16 tensors and is byte-identical to the completed v3 checkpoint at update 30,000.
SHA-256: 0b42edee4a2367e0b0aa51746778cc34cfc8c50716cce5ba7db2237533dc4323.
Modification notice: this separately trained LoRA changes Qwen Image 2.1 transformer attention projections when loaded. Distribution and use are subject to the Qwen Research License Agreement. See Notice for upstream attribution.
Model tree for SimpleTuner/Qwen-Image-2.1-training-assistant-v3
Base model
Qwen/Qwen-Image-2.1














































