Download docs/DETAILS.md from Haverbex/Qwen-Image-2.1-UltraFast: direct link, hf CLI and curl.
- Browser
- Download file 5.99 kB
-
https://huggingface.co/Haverbex/Qwen-Image-2.1-UltraFast/resolve/main/docs/DETAILS.md
- Command line
-
hf download hf://Haverbex/Qwen-Image-2.1-UltraFast/docs/DETAILS.md
-
curl -L -o DETAILS.md https://huggingface.co/Haverbex/Qwen-Image-2.1-UltraFast/resolve/main/docs/DETAILS.md
Model details and experimental results
← Model card and comparison gallery · Setup and usage
This page contains the training configuration, checkpoint provenance, full benchmark charts, and validation limits for Qwen-Image-2.1-UltraFast. The banner on the main page is branding artwork, not a model-output example.
Checkpoint details
| Item | Value |
|---|---|
| Base model | Qwen/Qwen-Image-2.1 |
| Base revision | 790c92633540aa0cb11d9abf19eb46d861714758 |
| Adapter | Custom LoRA, rank 32, alpha 32, BF16 |
| Target modules | 224 transformer projections |
| Trainable parameters | 83,886,080 |
| Training inputs | 16 MagicBrush train cases, English/local edits |
| Teacher targets | Eight intervals from each 40-step trajectory; 128 total |
| Optimizer | SGD, learning rate 0.1, momentum 0.9 |
| Updates | 128 |
| Tested output | 1024×1024 RGB, batch1, one reference image |
| Adapter SHA-256 | 9d3585e6f1b4700130a12df83b2ac7a9a7440158bc9625425475c28214e59639 |
The frozen base weights did not change. The same-runtime fresh-process replay compared 899 tensors plus optimizer/RNG/cursor and update records, with exact equality. This does not establish reproducibility across different hardware or CUDA builds.
Benchmark charts
Charts are compiled with Microsoft Flint (flint-chart@0.5.1, Vega-Lite backend). The underlying measurement data, Flint specifications, and render verification are included.
Recorded denoising time
| Measurement | Original T40 | UltraFast S8 |
|---|---|---|
| Expression edit, seconds | 69.69 | 18.08 |
| Flower sticker, seconds | 68.77 | 17.69 |
| Mean, seconds | 69.23 | 17.89 |
| Denoising model calls | 40 | 8 |
| Image-quality retention | Reference; no matched quality score | Not measured |
The historical mean time ratio is 3.87×, or 74.16% less denoising-pipeline time. These are two training cases measured in separate L4/BF16/1024 runs. Loading, encoders and VAE decoding are excluded; T40 additionally records trajectories. This is not a controlled end-to-end acceleration benchmark, and no quality-retention percentage has been measured.
Image-quality retention
Quality retention is not measured. The original model is the reference, but neither version has a matched human or independent editing-quality score in this study. Training MSE and the two visual examples are separate diagnostics. The Flint status chart therefore shows the missing measurement without assigning either model a quality percentage.
Training loss and interval regressions
Mean training velocity-MSE changed from 0.02848303 to 0.01230226 (−56.81%) across 128 paired intervals. It is a training objective, not an image-quality score.
Positive values indicate worse loss. Interval groups 1–4 worsened; across individual intervals, 63 improved and 65 worsened. These plots retain those regressions rather than presenting the overall mean as uniform improvement.
Observed results and limitations
The same 128 training intervals had mean velocity-MSE 0.0284830323 before training and 0.0123022638 after training, a 56.81% reduction. All 16 case-level means improved. However, 63 individual intervals improved and 65 worsened. Interval groups 1–4 worsened by 25.97–39.45%; interval 0 accounts for 90.63% of net improvement. This is not improvement across the entire trajectory. Full interval/case summary · All 128 paired measurements.
Two fresh-process S8 rollouts completed eight model calls each and finite RGB decoding. Their 18.08s/17.69s pipeline times exclude model loading, input encoding, and VAE decoding. Matching older T40 records are 69.69s/68.77s. The historical ratio is approximately 3.87×, or 74.16% less denoising-pipeline time. This is not a controlled or certified speedup: the T40 timing includes trajectory-recording callbacks, the runs were separate, and there is no paired untrained 8-step control for these two cases. Both used L4/BF16/1024 and KV cache. The reduction also includes the 40→8 step-count change; it cannot be attributed solely to distillation.
The sample edits are qualitative checks on training data. There is no independent human evaluation, development-set success rate, established quality-retention percentage, four-step student, full Mix-STQ calibration, packed CUDA backend, or end-to-end latency qualification. This checkpoint should not be presented as a production-ready or lossless acceleration release.
License and attribution
The upstream Qwen Research License Agreement limits use to non-commercial research/evaluation unless a separate commercial license is obtained. Its derivative notices and naming requirements apply. This project is independent and is not an official Qwen release. This project is conducted for non-commercial research and evaluation. See NOTICE.
MagicBrush declares CC BY 4.0 for its dataset; source photographs in this experiment carry the individual declarations documented in ATTRIBUTION.md. Source-photo declarations and lineage were checked during dataset preparation; independent ownership and comprehensive redistribution/privacy clearance have not been certified. Generated edits remain accompanied by source attribution. The supplied gallery contains no publisher target images or masks.



