Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 10 days ago • 119
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning Paper • 2605.28774 • Published May 27 • 90
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR Paper • 2605.15726 • Published May 15 • 36
agent-distillation/Qwen2.5-32B-Instruct_cot_trajectories_2k Viewer • Updated Jun 9, 2025 • 3k • 48 • 1
agent-distillation/Qwen2.5-32B-Instruct_cot_trajectories_2k Viewer • Updated Jun 9, 2025 • 3k • 48 • 1