Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 6 days ago • 58
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Paper • 2605.15565 • Published May 15 • 16