Instructions to use PuppetVision/sdxl_base_and_refiner_01_int8_comfyui with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use PuppetVision/sdxl_base_and_refiner_01_int8_comfyui with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- β‘ SDXL Base + Refiner INT8 Tensorwise for ComfyUI
β‘ SDXL Base + Refiner INT8 Tensorwise for ComfyUI
Two-stage SDXL Base + Refiner inference with native ComfyUI int8_tensorwise diffusion models, while keeping CLIP-L, CLIP-G and the VAE in their proven upstream formats.
π Benchmark headline: on the warm-run average (seeds 1β3), the INT8 + AMD Flash-Attention path completed the full 25-step Base + Refiner workflow in 16.821s, with 13.570s sampling at 0.543s/it. That is 18.9% less workflow time than the official Stability AI FP32 baseline and 19.9% less workflow time than the ComfyOrg BF16 baseline. Sampling time was reduced by 14.5% versus Stability AI FP32 and 15.7% versus ComfyOrg BF16.
Repository: https://huggingface.co/PuppetVision/sdxl_base_and_refiner_01_int8_comfyui
Ready-to-use workflows: https://huggingface.co/PuppetVision/sdxl_base_and_refiner_01_int8_comfyui/tree/main/workflows
π¬ Performance benchmark video
π Benchmarks first
Every row in benchmarks/benchmarks.csv uses the same 25-step SDXL Base + Refiner workflow and the same seed sequence for the compared model/runtime combinations.
- Seed 0 β first workflow run / cold run.
- Seeds 1β3 β warm workflow runs at 1216Γ832; the table reports their average.
- Seed 4 β resolution-change workflow run at 1024Γ1024.
- PyTorch cross-attention rows were run with stock ComfyUI.
- Flash-Attention rows were run with the AMD-optimized ComfyUI environment, using the AMD Triton/AITER Flash-Attention path.
- All rows use 25 steps, no LoRA, the same FP16 CLIP-L/CLIP-G pair, and the same SDXL VAE configuration recorded in the CSV.
- The Stability AI baseline uses the official FP32 SDXL Base + Refiner model weights.
- The ComfyOrg baseline uses BF16 SDXL Base + Refiner model weights.
| Configuration | Seed 0 workflow | Seeds 1β3 workflow avg | Seeds 1β3 sampling avg | Warm sec/it | Seed 4 resolution-change workflow |
|---|---|---|---|---|---|
| Official Stability AI FP32 Β· PyTorch cross-attention | 26.835s | 20.746s | 15.866s | 0.635s | 21.003s |
| This release INT8 Tensorwise Β· PyTorch cross-attention | 25.264s | 18.167s | 14.397s | 0.576s | 18.766s |
| This release INT8 Tensorwise Β· Flash-Attention / Triton / AITER | 23.208s | 16.821s | 13.570s | 0.543s | 17.842s |
| Official ComfyOrg BF16 Β· PyTorch cross-attention | 37.207s | 21.009s | 16.102s | 0.644s | 21.288s |
What the numbers show
- INT8 alone matters: with stock ComfyUI + PyTorch cross-attention, warm workflow time falls from 20.746s with Stability AI FP32 to 18.167s with this INT8 release β 12.4% less workflow time. Warm sampling is 9.3% lower.
- AMD Flash-Attention compounds the gain: using the same INT8 models, the optimized AMD path cuts another 7.4% from warm workflow time and 5.7% from warm sampling versus the stock PyTorch-cross-attention INT8 run.
- Versus ComfyOrg BF16: the INT8 + AMD Flash-Attention path reduces warm workflow time by 19.9% and warm sampling time by 15.7%.
- Fastest measured warm result: 16.821s workflow, 13.570s sampling, 0.543s/it.
- The complete raw measurements, startup flags, environment capture, VAE timings, resolutions and per-seed values are retained in
benchmarks/benchmarks.csv.
Benchmark results are specific to the recorded hardware/software environment and should be treated as measured reference data, not a universal performance guarantee.
πΌοΈ Example output
π Runtime recommendations
NVIDIA β use PyTorch cross-attention
For NVIDIA systems, use stock ComfyUI with the PyTorch cross-attention path:
python main.py --use-pytorch-cross-attention
The PyTorch-cross-attention benchmark rows in this repository are the stock-ComfyUI comparison path.
AMD ROCm β build and use Flash-Attention with Triton/AITER
The AMD benchmark path uses ROCm PyTorch plus Flash-Attention's AMD Triton backend.
1. Start from a working ROCm PyTorch environment
Install a PyTorch build that matches your ROCm stack, activate that environment, then install the build prerequisites:
python -m pip install -U pip setuptools wheel packaging ninja
PyTorch ROCm installation selector: https://pytorch.org/get-started/locally/
2. Build Flash-Attention with the AMD Triton/AITER backend
git clone --recursive https://github.com/Dao-AILab/flash-attention.git
cd flash-attention
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
python -m pip install --no-build-isolation .
If you maintain a compatible system AITER installation, you can build against it instead:
git clone --recursive https://github.com/ROCm/aiter.git
cd aiter
# Keep the Triton installation already paired with your ROCm/PyTorch stack.
AITER_USE_SYSTEM_TRITON=1 python setup.py develop
cd ../flash-attention
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
export FLASH_ATTENTION_USE_SYSTEM_AITER=TRUE
python -m pip install --no-build-isolation .
AITER support is GPU- and operator-dependent. Validate your exact ROCm, PyTorch, Triton, AITER and GPU combination before relying on it in production.
3. Select the AMD Triton backend at runtime
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
Then launch ComfyUI with:
python main.py \
--use-flash-attention \
--disable-xformers
Optional Flash-Attention autotuning can improve some workloads at the cost of one-time warmup/compilation work:
export FLASH_ATTENTION_TRITON_AMD_AUTOTUNE=TRUE
Upstream references:
π§ Technical specifications
Diffusion models
| Component | Release format | Upstream architecture | Precision | Conditioning | Key architecture details |
|---|---|---|---|---|---|
| SDXL Base | Standalone ComfyUI diffusion model | UNet2DConditionModel |
Mixed FP16 + native int8_tensorwise Linear weights |
2048-d cross-attention | channels 320/640/1280; transformer depth 1/2/10; latent I/O 4 channels |
| SDXL Refiner | Standalone ComfyUI diffusion model | UNet2DConditionModel |
Mixed FP16 + native int8_tensorwise Linear weights |
1280-d cross-attention | channels 384/768/1536/1536; transformer depth 4; latent I/O 4 channels |
The quantized diffusion weights use ComfyUI's native tensorwise INT8 representation:
- eligible Linear weights are stored as signed INT8;
- each quantized weight has a per-output-row FP32
weight_scale; - tensors outside the selected quantization scope remain floating point;
- the files contain quantization metadata understood by current ComfyUI mixed-precision loading.
These are diffusion-model-only files intended for ComfyUI's Load Diffusion Model path, not monolithic checkpoints.
Text encoders β FP16, intentionally not quantized
| Encoder | Role | Precision | Hidden size | Layers | Attention heads |
|---|---|---|---|---|---|
| CLIP-L | SDXL Base encoder | FP16 | 768 | 12 | 12 |
| CLIP-G / OpenCLIP bigG | SDXL Base + Refiner encoder | FP16 | 1280 | 32 | 20 |
SDXL Base uses the CLIP-L + CLIP-G conditioning stack. SDXL Refiner uses the 1280-dimensional CLIP-G conditioning path.
The text encoders distributed with this repository are verified upstream FP16 artifacts and are not INT8-converted.
VAE
- File:
vae/sdxl_vae.safetensors - The VAE remains separate from the diffusion models and text encoders.
- Exact SHA-256 and upstream provenance are recorded in
provenance/manifest.json. - See
PROVENANCE.mdandLICENSES.mdfor the exact upstream source and licensing record.
π¦ ComfyUI file placement
ComfyUI/models/
βββ diffusion_models/
β βββ sdxl_base_1.0_int8_tensorwise.safetensors
β βββ sdxl_refiner_1.0_int8_tensorwise.safetensors
βββ text_encoders/
β βββ clip_l_sdxl_fp16.safetensors
β βββ clip_g_sdxl_fp16.safetensors
βββ vae/
βββ sdxl_vae.safetensors
Use the Base model with SDXL dual-encoder conditioning. Use the Refiner with the CLIP-G / SDXL Refiner conditioning path.
Ready-to-use ComfyUI workflow files are available under workflows/.
π Credits
- Stability AI β creators of Stable Diffusion XL Base 1.0 and SDXL Refiner 1.0, and the upstream SDXL architecture and weights used to create the diffusion-model derivatives in this repository.
https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0
https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0 - madebyollin β for the widely used SDXL FP16-safe VAE work and
sdxl-vae-fp16-fixrelease.
https://huggingface.co/madebyollin/sdxl-vae-fp16-fix - ComfyUI / Comfy-Org β native model loading and
int8_tensorwiseexecution infrastructure. - ROCm / AITER and Dao-AILab Flash-Attention β AMD Triton/AITER and Flash-Attention infrastructure used by the optimized AMD benchmark path.
See PROVENANCE.md, LICENSES.md, THIRD_PARTY_NOTICES.md, and provenance/manifest.json for exact artifact provenance and licensing records.
βοΈ Warranty and liability disclaimer
This repository and its files are provided "AS IS", without warranties or conditions of any kind, express or implied, including but not limited to merchantability, fitness for a particular purpose, non-infringement, availability, accuracy, performance, or compatibility with any specific hardware/software stack.
GPU kernels, quantized inference, ROCm/CUDA environments, third-party extensions and model execution can fail, produce incorrect output, or cause data loss or system instability. You are responsible for validating the files, commands, licenses, outputs and operational suitability for your own use.
To the maximum extent permitted by applicable law, the repository authors/contributors are not liable for direct, indirect, incidental, special, consequential or other damages arising from use of, inability to use, or reliance on this repository.
Upstream components remain subject to their own licenses and terms. This section is informational and does not replace the controlling license texts in licenses/.
πΌ Follow / Contact / More Projects
I am seeking AI Systems Engineering opportunities in Amsterdam involving model optimization, inference, GPU acceleration, quantization, ROCm/CUDA, Triton, and generative-AI infrastructure.
LinkedIn
https://www.linkedin.com/in/allen-b-3a35505a/
YouTube β PuppetVisionAI
https://www.youtube.com/@PuppetVisionAI
Website
https://puppetvision.nl
GitHub / More Projects
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton#work-with-me--more-projects
- Downloads last month
- 27
