Text-to-Image
lora
simpletuner
training-assistant
qwen-image
experimental

SimpleTuner Training Assistant v3 for Qwen Image 2.1

Built with Qwen. This training assistant LoRA was trained from scratch for 30,000 updates, with rank/alpha 64, four base resolutions, and flow schedule auto shift. It uses corrected adamw_bf16, learning rate 1e-4, and global gradient norm clipping at 1.0.

Keep this adapter frozen and active during downstream LoRA training, then disable it when generating images with the downstream adapter. Its purpose is to help preserve coherence and image quality during fine-tuning.

Compared with v2, v3 increases rank from 32 to 64 and assistant training from 1,000 to 30,000 updates, uses a larger data pool and effective batch size, and enables automatic flow-schedule shifting. V3 starts from a fresh adapter; it does not resume or expand v2. V1 used synthetic targets only; v2 and v3 also use real images.

Use in SimpleTuner

Use a SimpleTuner checkout with Qwen Image 2.1 and assistant-LoRA support. Add these fields to a complete concept-training configuration:

{
  "model_family": "qwen_image",
  "model_flavour": "v2.1",
  "model_type": "lora",
  "lora_type": "standard",
  "assistant_lora_path": "SimpleTuner/Qwen-Image-2.1-training-assistant-v3",
  "disable_assistant_lora": false,
  "assistant_lora_strength": 1.0,
  "assistant_lora_inference_strength": 0.0
}

SimpleTuner freezes the assistant alongside the trainable concept adapter. Use ordinary concept training; distillation_method: assistant_lora creates an assistant using online teacher generation. At inference, load the base model and your resulting concept LoRA with this assistant disabled. This assistant has no trigger word. It is not a character adapter or a few-step generation adapter. Older Qwen Image flavours have not been tested with these weights.

The assistant is intentionally trained to absorb changes that would otherwise affect downstream tuning. Images made with the assistant enabled by itself do not represent the intended final inference setup. The relevant evaluation is downstream training with the assistant enabled and frozen, followed by inference with it disabled.

Training

The source pool contains 30,000 unique images, divided equally among:

Each source group has one third of the sampling probability. Synthetic images retain their generated dimensions in 12 aspect/resolution backends. The same real-image IDs are reused across four resolution backends per source; these are not additional unique images. All eight ranks' completed caches were audited against these counts.

Setting Value
Hardware 8 ร— NVIDIA H100
Updates 30,000
Batch per process / accumulation / effective batch 2 / 1 / 16
Precision BF16
LoRA rank / alpha 64 / 64
Attention projections to_q, to_k, to_v, to_out.0
Trainable parameters 67,108,864
Optimizer Corrected adamw_bf16
Learning rate 1e-4, constant after 25 warmup steps
Gradient clipping Global norm, maximum 1.0
Gradient checkpointing Enabled, interval 1
Base resolutions 512, 1024, 1536, 2048; aspect buckets
Flow schedule Auto shift enabled
REPA / regularisation Neither used
VAE Original Qwen Image 2.1; tiling disabled, encoding batch 1

The resumed phase after checkpoint 1,000 recorded 29,000 sampled batches: 9,587 synthetic, 9,676 CC12M and 9,737 e621. Data are sampled repeatedly over 30,000 updates. The nominal 512x512 value in the weights metadata is not the complete list of training sizes.

See training_details.json for artifact provenance and resume notes, training_config.json and dataloader.json for the recipe, and prompts.json for validation prompts. The published config resets the resume setting for a fresh run.

Final assistant validation

The comparisons below show the bare base model and base + v3 assistant at update 30,000, with assistant strength 1.0, 40 inference steps, CFG 1 and validation seed 42. Columns are output resolutions. Original comparison grids were split into separate, full-resolution panels without resizing; the bottom prompt strip and dividing gutter were removed, while the embedded model labels were retained.

These show what the assistant itself changes. They do not show a downstream concept LoRA trained with v3 and then evaluated with the assistant disabled. V3's downstream benefit and its comparison with v2 remain to be evaluated. Existing LoRA experiments establish the context for earlier assistants, not a v3 result.

Samples use the original Qwen Image 2.1 VAE. Texture differences from Ollin's texture-fixed decoder are separate from the assistant's effect.

Expand: portrait โ€” base versus assistant v3 at 30k

A photograph of an elderly woman wearing a blue silk scarf, soft window light, a plain gray background.

Model 512px 1024px 1536px 2048px
Base portrait, base, 512px portrait, base, 1024px portrait, base, 1536px portrait, base, 2048px
Base + v3, 30k portrait, assistant-v3-30000, 512px portrait, assistant-v3-30000, 1024px portrait, assistant-v3-30000, 1536px portrait, assistant-v3-30000, 2048px
Expand: street โ€” base versus assistant v3 at 30k

A photograph of a woman walking down a rain-soaked city street, reflections in the pavement, natural evening light.

Model 512px 1024px 1536px 2048px
Base street, base, 512px street, base, 1024px street, base, 1536px street, base, 2048px
Base + v3, 30k street, assistant-v3-30000, 512px street, assistant-v3-30000, 1024px street, assistant-v3-30000, 1536px street, assistant-v3-30000, 2048px
Expand: interior โ€” base versus assistant v3 at 30k

A photograph of a sunlit kitchen with wooden cabinets, ceramic cups and fresh flowers on a table.

Model 512px 1024px 1536px 2048px
Base interior, base, 512px interior, base, 1024px interior, base, 1536px interior, base, 2048px
Base + v3, 30k interior, assistant-v3-30000, 512px interior, assistant-v3-30000, 1024px interior, assistant-v3-30000, 1536px interior, assistant-v3-30000, 2048px
Expand: landscape โ€” base versus assistant v3 at 30k

A photograph of a mountain lake at dawn with mist over the water and pine trees along the shore.

Model 512px 1024px 1536px 2048px
Base landscape, base, 512px landscape, base, 1024px landscape, base, 1536px landscape, base, 2048px
Base + v3, 30k landscape, assistant-v3-30000, 512px landscape, assistant-v3-30000, 1024px landscape, assistant-v3-30000, 1536px landscape, assistant-v3-30000, 2048px
Expand: fabric โ€” base versus assistant v3 at 30k

A close-up photograph of a white cotton shirt and a folded blue silk scarf on a wooden table.

Model 512px 1024px 1536px 2048px
Base fabric, base, 512px fabric, base, 1024px fabric, base, 1536px fabric, base, 2048px
Base + v3, 30k fabric, assistant-v3-30000, 512px fabric, assistant-v3-30000, 1024px fabric, assistant-v3-30000, 1536px fabric, assistant-v3-30000, 2048px
Expand: fox โ€” base versus assistant v3 at 30k

An illustration of an anthropomorphic fox wearing a blue jacket and reading a book in a cozy library.

Model 512px 1024px 1536px 2048px
Base fox, base, 512px fox, base, 1024px fox, base, 1536px fox, base, 2048px
Base + v3, 30k fox, assistant-v3-30000, 512px fox, assistant-v3-30000, 1024px fox, assistant-v3-30000, 1536px fox, assistant-v3-30000, 2048px

Artifact and license

pytorch_lora_weights.safetensors contains 256 finite BF16 tensors and is byte-identical to the completed v3 checkpoint at update 30,000.

SHA-256: 0b42edee4a2367e0b0aa51746778cc34cfc8c50716cce5ba7db2237533dc4323.

Modification notice: this separately trained LoRA changes Qwen Image 2.1 transformer attention projections when loaded. Distribution and use are subject to the Qwen Research License Agreement. See Notice for upstream attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for SimpleTuner/Qwen-Image-2.1-training-assistant-v3

Adapter
(98)
this model

Datasets used to train SimpleTuner/Qwen-Image-2.1-training-assistant-v3