⚑ SDXL Base + Refiner INT8 Tensorwise for ComfyUI

Two-stage SDXL Base + Refiner inference with native ComfyUI int8_tensorwise diffusion models, while keeping CLIP-L, CLIP-G and the VAE in their proven upstream formats.

πŸ† Benchmark headline: on the warm-run average (seeds 1–3), the INT8 + AMD Flash-Attention path completed the full 25-step Base + Refiner workflow in 16.821s, with 13.570s sampling at 0.543s/it. That is 18.9% less workflow time than the official Stability AI FP32 baseline and 19.9% less workflow time than the ComfyOrg BF16 baseline. Sampling time was reduced by 14.5% versus Stability AI FP32 and 15.7% versus ComfyOrg BF16.

Repository: https://huggingface.co/PuppetVision/sdxl_base_and_refiner_01_int8_comfyui
Ready-to-use workflows: https://huggingface.co/PuppetVision/sdxl_base_and_refiner_01_int8_comfyui/tree/main/workflows

🎬 Performance benchmark video

YouTube fallback preview

Watch the SDXL INT8 Tensorwise benchmark video

🏁 Benchmarks first

Every row in benchmarks/benchmarks.csv uses the same 25-step SDXL Base + Refiner workflow and the same seed sequence for the compared model/runtime combinations.

  • Seed 0 β€” first workflow run / cold run.
  • Seeds 1–3 β€” warm workflow runs at 1216Γ—832; the table reports their average.
  • Seed 4 β€” resolution-change workflow run at 1024Γ—1024.
  • PyTorch cross-attention rows were run with stock ComfyUI.
  • Flash-Attention rows were run with the AMD-optimized ComfyUI environment, using the AMD Triton/AITER Flash-Attention path.
  • All rows use 25 steps, no LoRA, the same FP16 CLIP-L/CLIP-G pair, and the same SDXL VAE configuration recorded in the CSV.
  • The Stability AI baseline uses the official FP32 SDXL Base + Refiner model weights.
  • The ComfyOrg baseline uses BF16 SDXL Base + Refiner model weights.
Configuration Seed 0 workflow Seeds 1–3 workflow avg Seeds 1–3 sampling avg Warm sec/it Seed 4 resolution-change workflow
Official Stability AI FP32 Β· PyTorch cross-attention 26.835s 20.746s 15.866s 0.635s 21.003s
This release INT8 Tensorwise Β· PyTorch cross-attention 25.264s 18.167s 14.397s 0.576s 18.766s
This release INT8 Tensorwise Β· Flash-Attention / Triton / AITER 23.208s 16.821s 13.570s 0.543s 17.842s
Official ComfyOrg BF16 Β· PyTorch cross-attention 37.207s 21.009s 16.102s 0.644s 21.288s

What the numbers show

  • INT8 alone matters: with stock ComfyUI + PyTorch cross-attention, warm workflow time falls from 20.746s with Stability AI FP32 to 18.167s with this INT8 release β€” 12.4% less workflow time. Warm sampling is 9.3% lower.
  • AMD Flash-Attention compounds the gain: using the same INT8 models, the optimized AMD path cuts another 7.4% from warm workflow time and 5.7% from warm sampling versus the stock PyTorch-cross-attention INT8 run.
  • Versus ComfyOrg BF16: the INT8 + AMD Flash-Attention path reduces warm workflow time by 19.9% and warm sampling time by 15.7%.
  • Fastest measured warm result: 16.821s workflow, 13.570s sampling, 0.543s/it.
  • The complete raw measurements, startup flags, environment capture, VAE timings, resolutions and per-seed values are retained in benchmarks/benchmarks.csv.

Benchmark results are specific to the recorded hardware/software environment and should be treated as measured reference data, not a universal performance guarantee.

πŸ–ΌοΈ Example output

SDXL Base + Refiner example output

πŸš€ Runtime recommendations

NVIDIA β€” use PyTorch cross-attention

For NVIDIA systems, use stock ComfyUI with the PyTorch cross-attention path:

python main.py --use-pytorch-cross-attention

The PyTorch-cross-attention benchmark rows in this repository are the stock-ComfyUI comparison path.

AMD ROCm β€” build and use Flash-Attention with Triton/AITER

The AMD benchmark path uses ROCm PyTorch plus Flash-Attention's AMD Triton backend.

1. Start from a working ROCm PyTorch environment

Install a PyTorch build that matches your ROCm stack, activate that environment, then install the build prerequisites:

python -m pip install -U pip setuptools wheel packaging ninja

PyTorch ROCm installation selector: https://pytorch.org/get-started/locally/

2. Build Flash-Attention with the AMD Triton/AITER backend

git clone --recursive https://github.com/Dao-AILab/flash-attention.git
cd flash-attention

export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
python -m pip install --no-build-isolation .

If you maintain a compatible system AITER installation, you can build against it instead:

git clone --recursive https://github.com/ROCm/aiter.git
cd aiter

# Keep the Triton installation already paired with your ROCm/PyTorch stack.
AITER_USE_SYSTEM_TRITON=1 python setup.py develop

cd ../flash-attention
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
export FLASH_ATTENTION_USE_SYSTEM_AITER=TRUE
python -m pip install --no-build-isolation .

AITER support is GPU- and operator-dependent. Validate your exact ROCm, PyTorch, Triton, AITER and GPU combination before relying on it in production.

3. Select the AMD Triton backend at runtime

export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE

Then launch ComfyUI with:

python main.py \
  --use-flash-attention \
  --disable-xformers

Optional Flash-Attention autotuning can improve some workloads at the cost of one-time warmup/compilation work:

export FLASH_ATTENTION_TRITON_AMD_AUTOTUNE=TRUE

Upstream references:

🧠 Technical specifications

Diffusion models

Component Release format Upstream architecture Precision Conditioning Key architecture details
SDXL Base Standalone ComfyUI diffusion model UNet2DConditionModel Mixed FP16 + native int8_tensorwise Linear weights 2048-d cross-attention channels 320/640/1280; transformer depth 1/2/10; latent I/O 4 channels
SDXL Refiner Standalone ComfyUI diffusion model UNet2DConditionModel Mixed FP16 + native int8_tensorwise Linear weights 1280-d cross-attention channels 384/768/1536/1536; transformer depth 4; latent I/O 4 channels

The quantized diffusion weights use ComfyUI's native tensorwise INT8 representation:

  • eligible Linear weights are stored as signed INT8;
  • each quantized weight has a per-output-row FP32 weight_scale;
  • tensors outside the selected quantization scope remain floating point;
  • the files contain quantization metadata understood by current ComfyUI mixed-precision loading.

These are diffusion-model-only files intended for ComfyUI's Load Diffusion Model path, not monolithic checkpoints.

Text encoders β€” FP16, intentionally not quantized

Encoder Role Precision Hidden size Layers Attention heads
CLIP-L SDXL Base encoder FP16 768 12 12
CLIP-G / OpenCLIP bigG SDXL Base + Refiner encoder FP16 1280 32 20

SDXL Base uses the CLIP-L + CLIP-G conditioning stack. SDXL Refiner uses the 1280-dimensional CLIP-G conditioning path.

The text encoders distributed with this repository are verified upstream FP16 artifacts and are not INT8-converted.

VAE

  • File: vae/sdxl_vae.safetensors
  • The VAE remains separate from the diffusion models and text encoders.
  • Exact SHA-256 and upstream provenance are recorded in provenance/manifest.json.
  • See PROVENANCE.md and LICENSES.md for the exact upstream source and licensing record.

πŸ“¦ ComfyUI file placement

ComfyUI/models/
β”œβ”€β”€ diffusion_models/
β”‚   β”œβ”€β”€ sdxl_base_1.0_int8_tensorwise.safetensors
β”‚   └── sdxl_refiner_1.0_int8_tensorwise.safetensors
β”œβ”€β”€ text_encoders/
β”‚   β”œβ”€β”€ clip_l_sdxl_fp16.safetensors
β”‚   └── clip_g_sdxl_fp16.safetensors
└── vae/
    └── sdxl_vae.safetensors

Use the Base model with SDXL dual-encoder conditioning. Use the Refiner with the CLIP-G / SDXL Refiner conditioning path.

Ready-to-use ComfyUI workflow files are available under workflows/.

πŸ™ Credits

See PROVENANCE.md, LICENSES.md, THIRD_PARTY_NOTICES.md, and provenance/manifest.json for exact artifact provenance and licensing records.

βš–οΈ Warranty and liability disclaimer

This repository and its files are provided "AS IS", without warranties or conditions of any kind, express or implied, including but not limited to merchantability, fitness for a particular purpose, non-infringement, availability, accuracy, performance, or compatibility with any specific hardware/software stack.

GPU kernels, quantized inference, ROCm/CUDA environments, third-party extensions and model execution can fail, produce incorrect output, or cause data loss or system instability. You are responsible for validating the files, commands, licenses, outputs and operational suitability for your own use.

To the maximum extent permitted by applicable law, the repository authors/contributors are not liable for direct, indirect, incidental, special, consequential or other damages arising from use of, inability to use, or reliance on this repository.

Upstream components remain subject to their own licenses and terms. This section is informational and does not replace the controlling license texts in licenses/.

πŸ’Ό Follow / Contact / More Projects

I am seeking AI Systems Engineering opportunities in Amsterdam involving model optimization, inference, GPU acceleration, quantization, ROCm/CUDA, Triton, and generative-AI infrastructure.

LinkedIn
https://www.linkedin.com/in/allen-b-3a35505a/

YouTube β€” PuppetVisionAI
https://www.youtube.com/@PuppetVisionAI

Website
https://puppetvision.nl

GitHub / More Projects
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton#work-with-me--more-projects

Downloads last month
27
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support