Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Paper • 2610.01428 • Published 1 day ago • 5
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions Paper • 2609.38593 • Published 4 days ago • 3
Personalized Image Generation with Reasoning and Reflection Paper • 2610.00737 • Published 3 days ago • 2
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models Paper • 2609.37533 • Published 4 days ago • 27
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing Paper • 2609.37362 • Published 4 days ago • 6
Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 3 days ago • 4
RPTune: Learned Context Curation for LLM Catalog Search Paper • 2610.00964 • Published 1 day ago • 11
Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 1 day ago • 19
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 4 days ago • 25
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 1 day ago • 66
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 4 days ago • 127
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement Paper • 2609.39045 • Published 3 days ago • 78
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 7 days ago • 30
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure Paper • 2609.30074 • Published 9 days ago • 2
Self-Play Search Distillation for Large Language Model Reasoning Paper • 2609.30936 • Published 8 days ago • 4
WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing Paper • 2608.18486 • Published 6 days ago • 4