DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 19 days ago • 224
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 25 days ago • 69
J.O.S.I.E-2 Collection The new J.O.S.I.E. version 2 family of personality-driven language models trained entirely on Apple Silicon • 22 items • Updated Aug 7 • 11
Metis Collection Metis persistent-memory model family based on Qwen3.5, including 4B, 9B, and 27B parameter scales. • 4 items • Updated Jul 31 • 8
Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. https://arxiv.org/abs/2609.00791 • 6 items • Updated 17 days ago • 19
Laguna M.1 Collection Our first M-class coding agent model, designed for long-horizon work. Apache 2.0. • 4 items • Updated Jul 21 • 24
Gemma 4 Collection Gemma 4 is Google's new model family including including E2B, E4B, 26B-A4B, and 31B. • 43 items • Updated Aug 26 • 267
Zamba2-VL Collection A suite of vision-language models based on Zamba2. • 3 items • Updated Jun 9 • 5