|
Download .venv/transformers/docs/source/en/quantization/spqr.md from DrDavis/PythonProject1: direct link, hf CLI and curl.
- Browser
- Download file 1.58 kB
-
https://huggingface.co/DrDavis/PythonProject1/resolve/main/.venv/transformers/docs/source/en/quantization/spqr.md
- Command line
-
hf download hf://DrDavis/PythonProject1/.venv/transformers/docs/source/en/quantization/spqr.md
-
curl -L -o spqr.md https://huggingface.co/DrDavis/PythonProject1/resolve/main/.venv/transformers/docs/source/en/quantization/spqr.md
1.58 kB
SpQR
SpQR quantization algorithm involves a 16x16 tiled bi-level group 3-bit quantization structure, with sparse outliers as detailed in SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.
To SpQR-quantize a model, refer to the Vahe1994/SpQR repository.
Load a pre-SpQR-quantized model in [~PreTrainedModel.from_pretrained].
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
quantized_model = AutoModelForCausalLM.from_pretrained(
"elvircrn/Llama-2-7b-SPQR-3Bit-16x16-red_pajama-hf",
torch_dtype=torch.half,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("elvircrn/Llama-2-7b-SPQR-3Bit-16x16-red_pajama-hf")