please create a guff from the e4m3fn version too

#7
by Gompa - opened
QuantStack org

It doesnt make sense to create ggufs from fp8 versions when bf16 or f32 is available, because gguf does the quantization better and so the error is smaller with the same file size. Since our quants are based on the original model the quality is superior compared to ones based on fp8.

YarvixPA changed discussion status to closed

don't the e4m3fn versions also render with only 4 steps instead of 50 ? that was the case for the old model

QuantStack org

that would be because they merged the lightning 4 step lora, you can just use that (;

oh the 4 step is the lora, but the e4m3fn version is 20 steps ? or am i misunderstanding the workflow notes? :

Model Steps CFG
Offical 50 4.0
fp8_e4m3fn 20 2.5
fp8_e4m3fn + 4steps LoRA 4 1.0
QuantStack org

if you dont use the 4 steps lora, you should use 20 steps

Sign up or log in to comment