view article Article KV Caching Explained: Optimizing Transformer Inference Efficiency not-lain • Jan 30, 2025 • 385
view article Article Prefill and Decode for Concurrent Requests - Optimizing LLM Performance tngtech • Apr 16, 2025 • 86
Qwen2.5 Collection Qwen2.5 language models, including pretrained and instruction-tuned models of 7 sizes, including 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B. • 43 items • Updated Mar 2 • 733
Leaderboards for Arabic Collection A collection for all leaderboards related to the Arabic Language. • 5 items • Updated Dec 9, 2025 • 8
view article Article Introducing smolagents: simple agents that write actions in code. +1 m-ric, merve, thomwolf • Dec 31, 2024 • 1.21k
view article Article Hugging Face x LangChain : A new partner package +1 Jofthomas, kkondratenko, efriis • May 14, 2024 • 162