view article Article Giving a 753B Model Eyes: Reproducing GLM-5.2 Vision on 4× 96 GB GPUs 0xSero • 4 days ago • 1
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 13 days ago • 184
Cerebras REAP Collection Sparse MoE models compressed using REAP (Router-weighted Expert Activation Pruning) method • 30 items • Updated Feb 25 • 152
view article Article From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels drbh, danieldk • Aug 18, 2025 • 109
WebWorld: A Large-Scale World Model for Web Agent Training Paper • 2602.14721 • Published Feb 16 • 19
ProgramBench: Can Language Models Rebuild Programs From Scratch? Paper • 2605.03546 • Published May 5 • 4
view article Article Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents nvidia • Apr 28 • 62
LeWM Collection Official checkpoints and datasets related to LeWM paper. • 9 items • Updated Mar 27 • 54
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration Paper • 2604.14116 • Published Apr 15 • 13