Chinese-Jev: Bringing System One Model to Chinese-Language Tasks Paper • 2609.36965 • Published 3 days ago • 15
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 16 days ago • 19
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 18 days ago • 51
Graph Machine: Towards Better Pretraining via Edges Paper • 2609.02881 • Published about 1 month ago • 7
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published 26 days ago • 9
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 29 days ago • 114
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 24 days ago • 72
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation Paper • 2609.03756 • Published 29 days ago • 25
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published Aug 27 • 40
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss Paper • 2609.00591 • Published Sep 1 • 18
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation Paper • 2608.24293 • Published Aug 25 • 14
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Paper • 2608.01049 • Published Aug 2 • 13
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Paper • 2608.00079 • Published Jul 29 • 19
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling Paper • 2606.24447 • Published Jun 23 • 1
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published Jul 2 • 69