When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis Paper • 2609.15309 • Published 18 days ago • 13
MInTRL: Off-policy Intervention can boost On-policy RL Paper • 2609.12419 • Published 21 days ago • 14
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 18 days ago • 33
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 18 days ago • 84
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 18 days ago • 250
Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing Paper • 2602.03845 • Published Feb 3 • 27
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Paper • 2509.09675 • Published Sep 11, 2025 • 28
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Paper • 2509.07980 • Published Sep 9, 2025 • 106
Learning to Reason via Mixture-of-Thought for Logical Reasoning Paper • 2505.15817 • Published May 21, 2025 • 18