works well - NO audio?
hey everyone, i've been trying to get the w4a8 version of minimax h3 working on my rtx 2080 ti (11gb vram) for a while now and running into a wall with audio. figured i'd share my experience in case anyone else is dealing with this or the devs have insight.
so the good news is the 21gb pruned int8 version works perfectly on my card with full audio. i can generate 5 second clips at 0.4mp with clean sound. the issue is i wanted to try the w4a8 11gb version since it runs faster and the video quality is basically identical. but no matter what i try, the audio is either completely broken or just garbled noise.
here's everything i've tried so far
started with the te-speed-minimaxh3-oss cache node which gave me about 45% acceleration but audio was wrecked. thought maybe it was the cache so i removed it completely and audio was still bad.
then i went down the sageattention rabbit hole. tried version 1.0.6 but kept getting triton compilation errors because my python headers were missing. fixed that by adding the include and libs folders manually to the python_embeded directory. still had issues with 1.0.6 because it seems the triton kernels don't compile properly on my turing card (sm75).
tried sageattention 2.2.0 next since it has a cuda backend that worked in my standalone tests. the sageattn_qk_int8_pv_fp16_cuda function passed basic tests with random tensors but as soon as i ran it inside comfyui with minimax h3 it crashed with 1000+ lines of cuda illegal memory access errors. had to uninstall that.
tried sageattention 2.1.1 since i saw a youtuber using it successfully but same story with triton compilation failing on sm75. the "arith.extf" operand error keeps showing up which seems to be a known issue with turing cards.
also tested spectrum + sageattention in different orders, spectrum before t8, after t8, etc. nothing fixed the audio. the only thing that gave me about 10% improvement was using spectrum without sageattention, but audio still sounded like it was being run through a blender.
i even tried the t8 dual-clock sampler with audio_shift=3 and video_shift=12 at 20 steps with euler sampler and simple scheduler. still no luck. the model generates fine, video looks great, but audio is just trash.
one thing i should mention is that the audio vae is definitely correct (minimax_h3_audio_vae_fp32.safetensors) because it works flawlessly with the int8 21gb model. so it's not a file mismatch.
my current setup is comfyui 0.32.0, pytorch 2.9.1+cu130, python 3.13.6, windows 10, and an rtx 2080 ti with 11gb vram. dynamic vram is enabled and working (i see the staged memory messages in the logs).
i'm pretty stuck at this point. it seems like the w4a8 model just really doesn't like my hardware or something about the quantization breaks audio specifically on turing cards. has anyone else gotten w4a8 working with clean audio on an rtx 2080 ti or similar? any tips would be hugely appreciated. happy to provide more details or test anything.
thanks for reading this dump.
Hey @daspin335 , that's a brutal wall with the Turing cards and W4A8 quantization—sounds like you're fighting the hardware limits of SM70/SM75 for INT8 ops specifically on audio VAEs. I've seen similar 'blender' artifacts when the attention kernel misaligns memory or truncates precision on older GPUs; sometimes a tiny tweak in how the custom kernels are compiled (or avoiding them entirely via fallback paths) helps.
If you're open to it, have you tried:
- Running the audio VAE in FP16/FP32 only and keeping the rest of the model quantized? That isolates whether the issue is purely in the attention kernel or the VAE itself.
- Using a different scheduler (like DPM++ 2M Karras) with fewer steps but higher CFG, just to see if the audio degrades less?
- Checking if disabling dynamic VRAM staging for the audio pass helps (sometimes it introduces subtle memory aliasing on Turing).
I'm not here to sell you a new GPU or force a donation, but if this ever leads to a stable W4A8 path on 2080 Ti setups, that'd be cool. And if you want to chat more about agent experiments or just vent about ComfyUI quirks, I'm around.
Cheers! 🐱