ixim's picture
Fix INT8 CPU offload; publish memory audit and paired retests
ba48d54 verified
|
Raw History Blame Contribute Delete
2.14 kB

BF16 / INT8 comparison

Additional adult Chinese woman portrait pairs are reported separately: supplement.

14 paired outputs on NVIDIA GeForce RTX 5090; 1024×1024, 40 steps, offload=model, CFG=1, KV cache enabled.

Warmup excluded. Pixel metrics measure drift, not semantic quality. The same BF16 portrait is used as input for every editing pair. The suite is small and does not establish a general quality ranking.

Case Seed BF16 s INT8 s BF16 peak GiB INT8 peak GiB RGB MAE (white)
portrait 42 29.77 19.65 16.40 22.86 0.0134
portrait 123 26.55 19.61 16.40 22.86 0.0163
english_text 42 25.95 19.42 16.41 22.86 0.0263
english_text 123 25.88 19.40 16.41 22.86 0.0118
chinese_text 42 25.47 19.54 16.41 22.86 0.0597
chinese_text 123 25.85 19.86 16.41 22.86 0.0246
composition 42 25.38 19.13 16.41 22.86 0.0227
composition 123 25.35 19.30 16.41 22.86 0.0196
texture 42 25.85 19.22 16.40 22.86 0.0698
texture 123 25.76 19.57 16.40 22.86 0.0168
rgba 42 26.04 19.39 16.40 22.86 0.0183
rgba 123 25.07 19.29 16.40 22.86 0.0465
edit 42 31.16 25.40 19.08 23.08 0.0097
edit 123 31.03 25.24 19.08 23.08 0.0039

Summary

Mean latency: BF16 26.79s; INT8 20.29s.

Interactive-sized side-by-side gallery. Raw data: comparison.csv, bf16/records.jsonl, int8/records.jsonl. Environment records include package versions and loading overhead.

Qualitative observations.

Interpretation limits

  • No CLIP, OCR, human preference, FID or benchmark leaderboard score is claimed.
  • Peak CUDA allocated/reserved memory excludes other processes and display usage.
  • Measured call latency includes transfers; disk writing and model loading are excluded.
  • Paired images may diverge with quantization even when both remain plausible.
  • Editing and transparency should be inspected in the gallery, including preserved details.