Instructions to use Vaelico/Wulver with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Vaelico/Wulver with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Vaelico/Wulver", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Are there any other quantization options / a LoRA-extracted version?
I have an 8 GB GPU, and as much as I'd like to use it, the speed with INT8 would be terrible because of the offloading, while the GGUF quantizations on Civitai aren't really worth the quality loss.
So basically, there are two ways around this: either release a LoRA-extracted version from the BASE model, or make a W4A8 ConvRot quantization (INT4 model weights + INT8 activations). The result is a quant that should have better quality than NVFP4. Official support was only recently added to ComfyUI (V32.0+ I think).
Anyway, I'd be insanely grateful if you could do either of these, because with my limited hardware I simply can't make the quantization / LoRA extraction myself.
Hi, it is up the W4A8 convrot model.