Are there any other quantization options / a LoRA-extracted version?

#2
by nosok12313 - opened

I have an 8 GB GPU, and as much as I'd like to use it, the speed with INT8 would be terrible because of the offloading, while the GGUF quantizations on Civitai aren't really worth the quality loss.

So basically, there are two ways around this: either release a LoRA-extracted version from the BASE model, or make a W4A8 ConvRot quantization (INT4 model weights + INT8 activations). The result is a quant that should have better quality than NVFP4. Official support was only recently added to ComfyUI (V32.0+ I think).

Anyway, I'd be insanely grateful if you could do either of these, because with my limited hardware I simply can't make the quantization / LoRA extraction myself.

Hi, it is up the W4A8 convrot model.

Sign up or log in to comment