Spaces:
Runtime error
Runtime error
|
Download diffusers/docs/source/en/api/quantization.md from theSure/Omnieraser: direct link, hf CLI and curl.
- Browser
- Download file 1.45 kB
-
https://huggingface.co/spaces/theSure/Omnieraser/resolve/main/diffusers/docs/source/en/api/quantization.md
- Command line
-
hf download hf://spaces/theSure/Omnieraser/diffusers/docs/source/en/api/quantization.md
-
curl -L -o quantization.md https://huggingface.co/spaces/theSure/Omnieraser/resolve/main/diffusers/docs/source/en/api/quantization.md
1.45 kB
A newer version of the Gradio SDK is available: 6.30.0
Quantization
Quantization techniques reduce memory and computational costs by representing weights and activations with lower-precision data types like 8-bit integers (int8). This enables loading larger models you normally wouldn't be able to fit into memory, and speeding up inference. Diffusers supports 8-bit and 4-bit quantization with bitsandbytes.
Quantization techniques that aren't supported in Transformers can be added with the [DiffusersQuantizer] class.
Learn how to quantize models in the Quantization guide.
BitsAndBytesConfig
[[autodoc]] BitsAndBytesConfig
GGUFQuantizationConfig
[[autodoc]] GGUFQuantizationConfig
TorchAoConfig
[[autodoc]] TorchAoConfig
DiffusersQuantizer
[[autodoc]] quantizers.base.DiffusersQuantizer