SimpleTuner Training Assistant v1 for Qwen Image 2.1

Built with Qwen. This experimental training assistant LoRA was trained from scratch for 1,000 updates, using corrected adamw_bf16, learning rate 1e-4 and global gradient norm clipping at 1.0. It replaces the earlier v1 weights in this repository.

Keep this adapter frozen and active during downstream LoRA training, then disable it when generating images with the downstream adapter. Read the evaluation below before choosing it for an experiment.

Use in SimpleTuner

Use a SimpleTuner checkout with Qwen Image 2.1 and assistant-LoRA support. Add these fields to a complete concept-training configuration:

{
  "model_family": "qwen_image",
  "model_flavour": "v2.1",
  "model_type": "lora",
  "lora_type": "standard",
  "assistant_lora_path": "SimpleTuner/Qwen-Image-2.1-training-assistant-v1",
  "assistant_lora_weight_name": "pytorch_lora_weights.safetensors",
  "disable_assistant_lora": false,
  "assistant_lora_strength": 1.0,
  "assistant_lora_inference_strength": 0.0
}

SimpleTuner freezes the assistant alongside the trainable concept adapter. Use ordinary concept training; distillation_method: assistant_lora creates an assistant using online teacher generation. At inference, load the base model and your resulting concept LoRA with this assistant disabled. This assistant has no trigger word. It is not a character adapter or a few-step generation adapter. Older Qwen Image flavours have not been tested with these weights.

Training

Training used the reusable set of 10,000 generated images, pinned to dataset revision 7fb60061955890fbcc697214cbfe9f9f5fd44ce9. Eight H100 GPUs generated the images from the unmodified Qwen Image 2.1 base model with 40 native inference steps, CFG 1 and BF16. Captions use CC12M's long_caption field; original CC12M images were not used as targets.

The images were decoded with the full-frame VAE and stored as PNGs. Training re-encodes those images, so the targets differ from directly cached terminal teacher latents. The L40S run used extracted copies of the published shards. The SimpleTuner example qwen_image-2.1-assistant-lora-offline.peft-lora reads the same data through Webshart.

Setting Value
GPU NVIDIA L40S
Updates / batch / accumulation 1,000 / 1 / 1
Precision BF16
LoRA rank / alpha 32 / 32
Attention projections to_q, to_k, to_v, to_out.0
Optimizer Corrected adamw_bf16
Learning rate 1e-4, constant after 25 warmup steps
Gradient clipping Global norm, maximum 1.0
Gradient checkpointing Enabled, interval 2
Sampling Equal weight for each of 12 image backends; no repeats
VAE Full-frame encode/decode; encoding batch 1
Trainable parameters 33,554,432

Training samples square, portrait and landscape buckets at four base resolutions:

Base resolution Square Portrait Landscape
512 512 × 512 384 × 672 672 × 384
1024 1024 × 1024 768 × 1344 1344 × 768
1536 1536 × 1536 1152 × 2016 2016 × 1152
2048 2048 × 2048 1536 × 2688 2688 × 1536

The run budgets 1,000 sampled images from the 10,000-image pool; it does not make a complete pass over that dataset. The nominal 1024x1024 field in the weights metadata is not the complete list of training sizes. See training_details.json for provenance.

Evaluation and limitations

This is an experimental assistant; successful Domokun concept training was not demonstrated.

The assistant itself was visually checked on four prompts at 1024 and 2048 pixels after 1,000 updates. The requested subjects remained recognizable, with composition and style changes; bicycle geometry artifacts remained. These checks do not establish broad quality preservation.

A downstream comparison trained fresh rank-32 concept adapters on RareConcepts/Domokun for 1,000 updates at 2048 pixels, using the brown-square emoji trigger (🟫). Both runs used corrected adamw_bf16, learning rate 1e-4, global norm clipping at 1.0, batch size 1, matching initial concept weights and initial training RNG states. Baseline validation images were pixel-identical. The control continued from an audited step-250 checkpoint to step 1,000 with unchanged data and batch settings. The assisted run used this frozen assistant at strength 1.0 during training and disabled it for every validation. Validation used 40 native inference steps and CFG 1.

At update 1,000:

  • Neither run reliably produced Domokun as the subject of the beach or red-scarf prompts; the main subjects were still people.
  • The control introduced strong Domokun-like features into an unrelated fox prompt: a brown rectangular creature with a red mouth and triangular teeth. The assisted run retained a recognizable orange fox reading a book, with a substantial shift toward cartoon illustration.
  • An unrelated elderly-woman portrait remained coherent in both runs.

This small comparison suggests reduced concept spillover on the tested fox prompt, but does not demonstrate a successful concept-learning recipe or a general downstream quality benefit. It covers one training seed, two character prompts and two unrelated prompts. Do not interpret preserved unrelated subjects alone as evidence that the new concept has been learned.

The optimizer's stochastic addition now computes input + alpha * other. The earlier v1 used the reversed coefficient placement, a learning rate of 1e-5, and individual-value clipping at 0.01. This release starts with fresh adapter and optimizer state. These changes do not by themselves establish a downstream quality benefit.

Artifact and license

pytorch_lora_weights.safetensors contains 256 finite BF16 tensors and matches the completed checkpoint at update 1,000.

SHA-256: 2571f7fa98824bee87b7fc2a8d6adf7347a9e9032e5b5fc5908e5ac76ce87f9d.

Modification notice: this separately trained LoRA changes Qwen Image 2.1 transformer attention projections when loaded. Distribution and use are subject to the Qwen Research License Agreement, including its non-commercial terms. See Notice for upstream attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SimpleTuner/Qwen-Image-2.1-training-assistant-v1

Adapter
(111)
this model

Dataset used to train SimpleTuner/Qwen-Image-2.1-training-assistant-v1