RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States Paper • 2609.12814 • Published 24 days ago • 1
Mahalanobis-Based Multi-Head Attention for Complex State Propagation Paper • 2608.24462 • Published Aug 25 • 1
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention Paper • 2609.24797 • Published 14 days ago • 11
Cross-Model Memory Transfer via Target-Side Reader Adaptation Paper • 2608.17050 • Published Aug 17 • 8
MemoryAthena: Adaptive Routing over Latent and Generated Memories Paper • 2609.25853 • Published 13 days ago • 15
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 21 days ago • 175
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards Paper • 2604.26733 • Published May 15 • 1
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 24 days ago • 266
Texture Generation on 3D Meshes with Point-UV Diffusion Paper • 2308.10490 • Published Aug 21, 2023 • 1
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression Paper • 2604.02119 • Published Apr 2 • 1
OASIS: Online Activation Subspace Learning for Memory-Efficient Training Paper • 2604.09406 • Published Apr 10 • 1
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 26 days ago • 45
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published about 1 month ago • 19
Negative Self-Distillation: Learning to Reason by Avoiding Flaws Paper • 2609.11699 • Published 25 days ago • 37
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 62
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention Paper • 2607.04422 • Published Aug 7 • 1