S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published Aug 31 • 148
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction Paper • 2610.01762 • Published 4 days ago • 166
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 24 days ago • 85
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 567
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 52
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 180
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents Paper • 2608.01827 • Published Aug 3 • 19