Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Paper • 2610.01428 • Published 3 days ago • 9
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions Paper • 2609.38593 • Published 5 days ago • 8
Personalized Image Generation with Reasoning and Reflection Paper • 2610.00737 • Published 4 days ago • 7
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models Paper • 2609.37533 • Published 5 days ago • 55
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing Paper • 2609.37362 • Published 5 days ago • 13
Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 4 days ago • 15
RPTune: Learned Context Curation for LLM Catalog Search Paper • 2610.00964 • Published 3 days ago • 20
Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 3 days ago • 40
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 5 days ago • 45
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 3 days ago • 78
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 5 days ago • 128
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement Paper • 2609.39045 • Published 4 days ago • 80
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 8 days ago • 30
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure Paper • 2609.30074 • Published 10 days ago • 2
Self-Play Search Distillation for Large Language Model Reasoning Paper • 2609.30936 • Published 9 days ago • 4
WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing Paper • 2608.18486 • Published 7 days ago • 4