Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated Aug 27
Fine-tuning/Fine-tuned Collection Fine-tuning generative AI is the process of taking a pre-trained base model and training it further on a smaller, specific dataset. • 335 items • Updated Aug 27
Importance Matrix (Imatrix) Collection The Importance Matrix (iMatrix) is a data-driven calibration method used in low-bit quantization for LLMs. • 330 items • Updated Aug 16
GPT-Generated Unified Format (GGUF) Collection GPT-Generated Unified Format (GGUF) is a single-file binary format used to store and run large language models efficiently on consumer hardware. • 706 items • Updated 21 days ago
Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated Aug 27
Mixture of Experts (MoE) Collection Mixture of Experts (MoE) is a machine learning architecture that splits a large neural network into smaller sub-networks called "experts." • 202 items • Updated Aug 16
GPT-Generated Unified Format (GGUF) Collection GPT-Generated Unified Format (GGUF) is a single-file binary format used to store and run large language models efficiently on consumer hardware. • 706 items • Updated 21 days ago
Automatic-Speech-Recognition (ASR) Collection Automatic Speech Recognition (ASR) converts spoken audio into text. In generative AI advanced ASR models act as the ears of large multimodal models. • 5 items • Updated Aug 14
nvidia/nemotron-speech-streaming-en-0.6b Automatic Speech Recognition • 0.6B • Updated Aug 5 • 381k • 630
Agentic Generative AI Collection Agentic generative AI refers to autonomous systems built on LLMs that can plan, use tools, and execute multi-step workflows. • 57 items • Updated Aug 11
Fine-tuning/Fine-tuned Collection Fine-tuning generative AI is the process of taking a pre-trained base model and training it further on a smaller, specific dataset. • 335 items • Updated Aug 27
Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated Aug 27