view article Article What a 61M-Parameter CTC Model Can and Cannot Learn mohammedaly22 • 19 days ago • 2
view article Article Small Language Models (SLM): A Comprehensive Overview jjokah • Feb 22, 2025 • 177
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published Aug 27 • 40
Efficient Memory Management for Large Language Model Serving with PagedAttention Paper • 2309.06180 • Published Sep 12, 2023 • 76
view article Article ArabicWeb24: Creating a High Quality Arabic Web-only Pre-training Dataset MayFarhat • Aug 8, 2024 • 12
view article Article makeMoE: Implement a Sparse Mixture of Experts Language Model from Scratch AviSoori1x • May 7, 2024 • 125
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Paper • 2506.01844 • Published Jun 2, 2025 • 168