AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance Paper • 2508.06944 • Published Aug 9, 2025 • 3