The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 12 days ago • 19
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 14 days ago • 50
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published 22 days ago • 9
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 25 days ago • 113
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 20 days ago • 71
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation Paper • 2609.03756 • Published 25 days ago • 25
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published Aug 27 • 40
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss Paper • 2609.00591 • Published 27 days ago • 18