CAST: Game Solvers as Turn-Level Teachers for LLM Agents Paper • 2607.25308 • Published 4 days ago • 36
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Paper • 2607.26784 • Published 3 days ago • 25
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published 2 days ago • 35
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published 2 days ago • 30
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published 2 days ago • 39
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 12 days ago • 59
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published 23 days ago • 87
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 19 days ago • 84
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 21 days ago • 71
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Paper • 2601.05808 • Published Jan 9 • 38
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 23 days ago • 76
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published about 1 month ago • 67
OpenThinker-Agent2 Collection OpenThinker-Agent2: agentic SFT/RL datasets and 8B/32B models (cold-start SFT, RL, and the OpenThinkerAgent-32B release). • 11 items • Updated Jun 11 • 10
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Paper • 2606.26058 • Published Jun 24 • 67
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 94
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents Paper • 2606.04703 • Published Jun 3 • 26
V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts Paper • 2603.10848 • Published Mar 11 • 17
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions Paper • 2605.27141 • Published May 26 • 20