topabaem's picture
Make model card concise and image-focused; move full details to linked pages
a99bedb verified
|
Raw History Blame Contribute Delete
5.99 kB

Model details and experimental results

← Model card and comparison gallery · Setup and usage

This page contains the training configuration, checkpoint provenance, full benchmark charts, and validation limits for Qwen-Image-2.1-UltraFast. The banner on the main page is branding artwork, not a model-output example.

Checkpoint details

Item Value
Base model Qwen/Qwen-Image-2.1
Base revision 790c92633540aa0cb11d9abf19eb46d861714758
Adapter Custom LoRA, rank 32, alpha 32, BF16
Target modules 224 transformer projections
Trainable parameters 83,886,080
Training inputs 16 MagicBrush train cases, English/local edits
Teacher targets Eight intervals from each 40-step trajectory; 128 total
Optimizer SGD, learning rate 0.1, momentum 0.9
Updates 128
Tested output 1024×1024 RGB, batch1, one reference image
Adapter SHA-256 9d3585e6f1b4700130a12df83b2ac7a9a7440158bc9625425475c28214e59639

The frozen base weights did not change. The same-runtime fresh-process replay compared 899 tensors plus optimizer/RNG/cursor and update records, with exact equality. This does not establish reproducibility across different hardware or CUDA builds.

Benchmark charts

Charts are compiled with Microsoft Flint (flint-chart@0.5.1, Vega-Lite backend). The underlying measurement data, Flint specifications, and render verification are included.

Recorded denoising time

Historical denoising timing comparison

Measurement Original T40 UltraFast S8
Expression edit, seconds 69.69 18.08
Flower sticker, seconds 68.77 17.69
Mean, seconds 69.23 17.89
Denoising model calls 40 8
Image-quality retention Reference; no matched quality score Not measured

The historical mean time ratio is 3.87×, or 74.16% less denoising-pipeline time. These are two training cases measured in separate L4/BF16/1024 runs. Loading, encoders and VAE decoding are excluded; T40 additionally records trajectories. This is not a controlled end-to-end acceleration benchmark, and no quality-retention percentage has been measured.

Image-quality retention

Quality-retention measurement status

Quality retention is not measured. The original model is the reference, but neither version has a matched human or independent editing-quality score in this study. Training MSE and the two visual examples are separate diagnostics. The Flint status chart therefore shows the missing measurement without assigning either model a quality percentage.

Training loss and interval regressions

Training velocity-MSE before and after

Mean training velocity-MSE changed from 0.02848303 to 0.01230226 (−56.81%) across 128 paired intervals. It is a training objective, not an image-quality score.

Loss change across eight interval groups

Positive values indicate worse loss. Interval groups 1–4 worsened; across individual intervals, 63 improved and 65 worsened. These plots retain those regressions rather than presenting the overall mean as uniform improvement.

Observed results and limitations

The same 128 training intervals had mean velocity-MSE 0.0284830323 before training and 0.0123022638 after training, a 56.81% reduction. All 16 case-level means improved. However, 63 individual intervals improved and 65 worsened. Interval groups 1–4 worsened by 25.97–39.45%; interval 0 accounts for 90.63% of net improvement. This is not improvement across the entire trajectory. Full interval/case summary · All 128 paired measurements.

Two fresh-process S8 rollouts completed eight model calls each and finite RGB decoding. Their 18.08s/17.69s pipeline times exclude model loading, input encoding, and VAE decoding. Matching older T40 records are 69.69s/68.77s. The historical ratio is approximately 3.87×, or 74.16% less denoising-pipeline time. This is not a controlled or certified speedup: the T40 timing includes trajectory-recording callbacks, the runs were separate, and there is no paired untrained 8-step control for these two cases. Both used L4/BF16/1024 and KV cache. The reduction also includes the 40→8 step-count change; it cannot be attributed solely to distillation.

The sample edits are qualitative checks on training data. There is no independent human evaluation, development-set success rate, established quality-retention percentage, four-step student, full Mix-STQ calibration, packed CUDA backend, or end-to-end latency qualification. This checkpoint should not be presented as a production-ready or lossless acceleration release.

License and attribution

The upstream Qwen Research License Agreement limits use to non-commercial research/evaluation unless a separate commercial license is obtained. Its derivative notices and naming requirements apply. This project is independent and is not an official Qwen release. This project is conducted for non-commercial research and evaluation. See NOTICE.

MagicBrush declares CC BY 4.0 for its dataset; source photographs in this experiment carry the individual declarations documented in ATTRIBUTION.md. Source-photo declarations and lineage were checked during dataset preparation; independent ownership and comprehensive redistribution/privacy clearance have not been certified. Generated edits remain accompanied by source attribution. The supplied gallery contains no publisher target images or masks.