Optimised AWQ Quants for high-throughput deployments of Gemma2! Compatible with Transformers, TGI & VLLM π€
AI & ML interests
Optimised quants for high-throughput deployments! Compatible with Transformers, TGI & vLLM π€
Optimised Quants for high-throughput deployments! Compatible with Transformers, TGI & VLLM π€
-
hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4
Text Generation β’ 410B β’ Updated β’ 1.75k β’ 36 -
hugging-quants/Meta-Llama-3.1-405B-Instruct-BNB-NF4
Text Generation β’ 423B β’ Updated β’ 29 β’ 5 -
hugging-quants/Meta-Llama-3.1-405B-Instruct-GPTQ-INT4
Text Generation β’ 410B β’ Updated β’ 137 β’ 16 -
hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4
Text Generation β’ 71B β’ Updated β’ 230k β’ 110
Llama.cpp compatible quants for Llama 3.2 3B and 1B Instruct models.
-
hugging-quants/Llama-3.2-3B-Instruct-Q8_0-GGUF
Text Generation β’ 3B β’ Updated β’ 1.18k β’ 53 -
hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF
Text Generation β’ 3B β’ Updated β’ 22.8k β’ 35 -
hugging-quants/Llama-3.2-1B-Instruct-Q8_0-GGUF
Text Generation β’ 1B β’ Updated β’ 285k β’ 51 -
hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF
Text Generation β’ 1B β’ Updated β’ 123k β’ 30
Optimised AWQ Quants for high-throughput deployments of Gemma2! Compatible with Transformers, TGI & VLLM π€
Llama.cpp compatible quants for Llama 3.2 3B and 1B Instruct models.
-
hugging-quants/Llama-3.2-3B-Instruct-Q8_0-GGUF
Text Generation β’ 3B β’ Updated β’ 1.18k β’ 53 -
hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF
Text Generation β’ 3B β’ Updated β’ 22.8k β’ 35 -
hugging-quants/Llama-3.2-1B-Instruct-Q8_0-GGUF
Text Generation β’ 1B β’ Updated β’ 285k β’ 51 -
hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF
Text Generation β’ 1B β’ Updated β’ 123k β’ 30
Optimised Quants for high-throughput deployments! Compatible with Transformers, TGI & VLLM π€
-
hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4
Text Generation β’ 410B β’ Updated β’ 1.75k β’ 36 -
hugging-quants/Meta-Llama-3.1-405B-Instruct-BNB-NF4
Text Generation β’ 423B β’ Updated β’ 29 β’ 5 -
hugging-quants/Meta-Llama-3.1-405B-Instruct-GPTQ-INT4
Text Generation β’ 410B β’ Updated β’ 137 β’ 16 -
hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4
Text Generation β’ 71B β’ Updated β’ 230k β’ 110