view article Article Unlocking Longer Generation with Key-Value Cache Quantization RaushanTurganbay • May 16, 2024 • 58
Nemotron 3 Embed Collection Open embedding models for enterprise RAG, agentic retrieval, code search, and agent memory. • 3 items • Updated 11 days ago • 31
view article Article Building Conversational AI: A Deep Dive into Voice Agent Architectures and Best Practices abdeljalilELmajjodi • Sep 2, 2025 • 24
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 20 days ago • 60
Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated-gguf-4bit Text Generation • 35B • Updated Jun 26 • 5.35k • 9