Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 8 days ago • 62
World Action Modeling with Progressive Visual Planning Paper • 2610.02508 • Published 6 days ago • 80
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization Paper • 2610.00906 • Published 6 days ago • 77
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training Paper • 2609.36659 • Published 8 days ago • 78
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 5 days ago • 86
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 263
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published Sep 1 • 36
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 144
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 92
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 126
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 105
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Paper • 2607.11683 • Published Jul 13 • 150
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Paper • 2605.30993 • Published May 29 • 55
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses Paper • 2606.02373 • Published Jun 1 • 59
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Paper • 2606.12191 • Published Jun 10 • 73
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution Paper • 2606.10917 • Published Jun 9 • 77
OCC-RAG: Optimal Cognitive Core for Faithful Question Answering Paper • 2606.00683 • Published May 30 • 102
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 49