# Model details and experimental results [← Model card and comparison gallery](../README.md) · [Setup and usage](USAGE.md) This page contains the training configuration, checkpoint provenance, full benchmark charts, and validation limits for Qwen-Image-2.1-UltraFast. The banner on the main page is branding artwork, not a model-output example. ## Checkpoint details | Item | Value | |---|---| | Base model | Qwen/Qwen-Image-2.1 | | Base revision | `790c92633540aa0cb11d9abf19eb46d861714758` | | Adapter | Custom LoRA, rank 32, alpha 32, BF16 | | Target modules | 224 transformer projections | | Trainable parameters | 83,886,080 | | Training inputs | 16 MagicBrush train cases, English/local edits | | Teacher targets | Eight intervals from each 40-step trajectory; 128 total | | Optimizer | SGD, learning rate 0.1, momentum 0.9 | | Updates | 128 | | Tested output | 1024×1024 RGB, batch1, one reference image | | Adapter SHA-256 | `9d3585e6f1b4700130a12df83b2ac7a9a7440158bc9625425475c28214e59639` | The frozen base weights did not change. The same-runtime fresh-process replay compared 899 tensors plus optimizer/RNG/cursor and update records, with exact equality. This does not establish reproducibility across different hardware or CUDA builds. ## Benchmark charts Charts are compiled with [Microsoft Flint](https://github.com/microsoft/flint-chart) (`flint-chart@0.5.1`, Vega-Lite backend). The underlying [measurement data](../reports/benchmark-data.json), [Flint specifications](https://huggingface.co/Haverbex/Qwen-Image-2.1-UltraFast/tree/main/charts), and [render verification](../reports/chart-verification.json) are included. ### Recorded denoising time ![Historical denoising timing comparison](../assets/benchmark-latency.png) | Measurement | Original T40 | UltraFast S8 | |---|---:|---:| | Expression edit, seconds | 69.69 | 18.08 | | Flower sticker, seconds | 68.77 | 17.69 | | Mean, seconds | 69.23 | 17.89 | | Denoising model calls | 40 | 8 | | Image-quality retention | Reference; no matched quality score | Not measured | The historical mean time ratio is **3.87×**, or **74.16% less denoising-pipeline time**. These are two training cases measured in separate L4/BF16/1024 runs. Loading, encoders and VAE decoding are excluded; T40 additionally records trajectories. This is not a controlled end-to-end acceleration benchmark, and no quality-retention percentage has been measured. ### Image-quality retention ![Quality-retention measurement status](../assets/benchmark-quality-status.png) **Quality retention is not measured.** The original model is the reference, but neither version has a matched human or independent editing-quality score in this study. Training MSE and the two visual examples are separate diagnostics. The [Flint status chart](../charts/benchmark-quality-status.flint.json) therefore shows the missing measurement without assigning either model a quality percentage. ### Training loss and interval regressions ![Training velocity-MSE before and after](../assets/benchmark-training-loss.png) Mean training velocity-MSE changed from **0.02848303 to 0.01230226 (−56.81%)** across 128 paired intervals. It is a training objective, not an image-quality score. ![Loss change across eight interval groups](../assets/benchmark-interval-change.png) Positive values indicate worse loss. Interval groups 1–4 worsened; across individual intervals, 63 improved and 65 worsened. These plots retain those regressions rather than presenting the overall mean as uniform improvement. ## Observed results and limitations The same 128 training intervals had mean velocity-MSE 0.0284830323 before training and 0.0123022638 after training, a 56.81% reduction. All 16 case-level means improved. **However, 63 individual intervals improved and 65 worsened. Interval groups 1–4 worsened by 25.97–39.45%; interval 0 accounts for 90.63% of net improvement.** This is not improvement across the entire trajectory. [Full interval/case summary](../reports/evaluation.json) · [All 128 paired measurements](../reports/paired-loss.csv). Two fresh-process S8 rollouts completed eight model calls each and finite RGB decoding. Their 18.08s/17.69s pipeline times exclude model loading, input encoding, and VAE decoding. Matching older T40 records are 69.69s/68.77s. The historical ratio is approximately 3.87×, or 74.16% less denoising-pipeline time. **This is not a controlled or certified speedup:** the T40 timing includes trajectory-recording callbacks, the runs were separate, and there is no paired untrained 8-step control for these two cases. Both used L4/BF16/1024 and KV cache. The reduction also includes the 40→8 step-count change; it cannot be attributed solely to distillation. The sample edits are qualitative checks on training data. There is no independent human evaluation, development-set success rate, established quality-retention percentage, four-step student, full Mix-STQ calibration, packed CUDA backend, or end-to-end latency qualification. This checkpoint should not be presented as a production-ready or lossless acceleration release. ## License and attribution The upstream [Qwen Research License Agreement](../LICENSE) limits use to non-commercial research/evaluation unless a separate commercial license is obtained. Its derivative notices and naming requirements apply. This project is independent and is not an official Qwen release. This project is conducted for non-commercial research and evaluation. See [NOTICE](../NOTICE). MagicBrush declares CC BY 4.0 for its dataset; source photographs in this experiment carry the individual declarations documented in [ATTRIBUTION.md](../ATTRIBUTION.md). Source-photo declarations and lineage were checked during dataset preparation; independent ownership and comprehensive redistribution/privacy clearance have not been certified. Generated edits remain accompanied by source attribution. The supplied gallery contains no publisher target images or masks.