CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 16 days ago • 138
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published Aug 31 • 148
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report Paper • 2608.15763 • Published Aug 22 • 54
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment Paper • 2601.18292 • Published Jan 26 • 12