topabaem's picture
Make model card concise and image-focused; move full details to linked pages
a99bedb verified
|
Raw History Blame Contribute Delete
5.99 kB
# Model details and experimental results
[← Model card and comparison gallery](../README.md) · [Setup and usage](USAGE.md)
This page contains the training configuration, checkpoint provenance, full benchmark charts, and validation limits for Qwen-Image-2.1-UltraFast. The banner on the main page is branding artwork, not a model-output example.
## Checkpoint details
| Item | Value |
|---|---|
| Base model | Qwen/Qwen-Image-2.1 |
| Base revision | `790c92633540aa0cb11d9abf19eb46d861714758` |
| Adapter | Custom LoRA, rank 32, alpha 32, BF16 |
| Target modules | 224 transformer projections |
| Trainable parameters | 83,886,080 |
| Training inputs | 16 MagicBrush train cases, English/local edits |
| Teacher targets | Eight intervals from each 40-step trajectory; 128 total |
| Optimizer | SGD, learning rate 0.1, momentum 0.9 |
| Updates | 128 |
| Tested output | 1024×1024 RGB, batch1, one reference image |
| Adapter SHA-256 | `9d3585e6f1b4700130a12df83b2ac7a9a7440158bc9625425475c28214e59639` |
The frozen base weights did not change. The same-runtime fresh-process replay compared 899 tensors plus optimizer/RNG/cursor and update records, with exact equality. This does not establish reproducibility across different hardware or CUDA builds.
## Benchmark charts
Charts are compiled with [Microsoft Flint](https://github.com/microsoft/flint-chart) (`flint-chart@0.5.1`, Vega-Lite backend). The underlying [measurement data](../reports/benchmark-data.json), [Flint specifications](https://huggingface.co/Haverbex/Qwen-Image-2.1-UltraFast/tree/main/charts), and [render verification](../reports/chart-verification.json) are included.
### Recorded denoising time
![Historical denoising timing comparison](../assets/benchmark-latency.png)
| Measurement | Original T40 | UltraFast S8 |
|---|---:|---:|
| Expression edit, seconds | 69.69 | 18.08 |
| Flower sticker, seconds | 68.77 | 17.69 |
| Mean, seconds | 69.23 | 17.89 |
| Denoising model calls | 40 | 8 |
| Image-quality retention | Reference; no matched quality score | Not measured |
The historical mean time ratio is **3.87×**, or **74.16% less denoising-pipeline time**. These are two training cases measured in separate L4/BF16/1024 runs. Loading, encoders and VAE decoding are excluded; T40 additionally records trajectories. This is not a controlled end-to-end acceleration benchmark, and no quality-retention percentage has been measured.
### Image-quality retention
![Quality-retention measurement status](../assets/benchmark-quality-status.png)
**Quality retention is not measured.** The original model is the reference, but neither version has a matched human or independent editing-quality score in this study. Training MSE and the two visual examples are separate diagnostics. The [Flint status chart](../charts/benchmark-quality-status.flint.json) therefore shows the missing measurement without assigning either model a quality percentage.
### Training loss and interval regressions
![Training velocity-MSE before and after](../assets/benchmark-training-loss.png)
Mean training velocity-MSE changed from **0.02848303 to 0.01230226 (−56.81%)** across 128 paired intervals. It is a training objective, not an image-quality score.
![Loss change across eight interval groups](../assets/benchmark-interval-change.png)
Positive values indicate worse loss. Interval groups 1–4 worsened; across individual intervals, 63 improved and 65 worsened. These plots retain those regressions rather than presenting the overall mean as uniform improvement.
## Observed results and limitations
The same 128 training intervals had mean velocity-MSE 0.0284830323 before training and 0.0123022638 after training, a 56.81% reduction. All 16 case-level means improved. **However, 63 individual intervals improved and 65 worsened. Interval groups 1–4 worsened by 25.97–39.45%; interval 0 accounts for 90.63% of net improvement.** This is not improvement across the entire trajectory. [Full interval/case summary](../reports/evaluation.json) · [All 128 paired measurements](../reports/paired-loss.csv).
Two fresh-process S8 rollouts completed eight model calls each and finite RGB decoding. Their 18.08s/17.69s pipeline times exclude model loading, input encoding, and VAE decoding. Matching older T40 records are 69.69s/68.77s. The historical ratio is approximately 3.87×, or 74.16% less denoising-pipeline time. **This is not a controlled or certified speedup:** the T40 timing includes trajectory-recording callbacks, the runs were separate, and there is no paired untrained 8-step control for these two cases. Both used L4/BF16/1024 and KV cache. The reduction also includes the 40→8 step-count change; it cannot be attributed solely to distillation.
The sample edits are qualitative checks on training data. There is no independent human evaluation, development-set success rate, established quality-retention percentage, four-step student, full Mix-STQ calibration, packed CUDA backend, or end-to-end latency qualification. This checkpoint should not be presented as a production-ready or lossless acceleration release.
## License and attribution
The upstream [Qwen Research License Agreement](../LICENSE) limits use to non-commercial research/evaluation unless a separate commercial license is obtained. Its derivative notices and naming requirements apply. This project is independent and is not an official Qwen release. This project is conducted for non-commercial research and evaluation. See [NOTICE](../NOTICE).
MagicBrush declares CC BY 4.0 for its dataset; source photographs in this experiment carry the individual declarations documented in [ATTRIBUTION.md](../ATTRIBUTION.md). Source-photo declarations and lineage were checked during dataset preparation; independent ownership and comprehensive redistribution/privacy clearance have not been certified. Generated edits remain accompanied by source attribution. The supplied gallery contains no publisher target images or masks.