Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 5 days ago • 85
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models Paper • 2609.37533 • Published 7 days ago • 61
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models Paper • 2609.37533 • Published 7 days ago • 61
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding Paper • 2609.32019 • Published 11 days ago • 53
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models Paper • 2609.37533 • Published 7 days ago • 61
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 9 days ago • 47
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published Sep 1 • 36
Running 44 T-Search: an open agentic retriever 🔎 44 Open agentic retriever for hard multi-step search
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation Paper • 2601.22813 • Published Jan 30 • 63
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines? Paper • 2602.14111 • Published Feb 15 • 57