Building to the Test: Coding Agents Deliver What You Check, Not What You Requested Paper • 2606.28430 • Published Jun 26 • 9
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity Paper • 2607.00152 • Published about 1 month ago • 9
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published 30 days ago • 9
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Paper • 2607.01071 • Published 30 days ago • 31
Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published 3 days ago • 28
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published 4 days ago • 28
CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents Paper • 2607.25431 • Published 3 days ago • 74
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 4 days ago • 87
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs Paper • 2607.25669 • Published 3 days ago • 8
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 5 days ago • 119
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Paper • 2607.24720 • Published 4 days ago • 25
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published 7 days ago • 57
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 4 days ago • 81
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making Paper • 2607.14277 • Published 16 days ago • 10
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 8 days ago • 24
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 9 days ago • 30
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 7 days ago • 45
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published 8 days ago • 8