Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Paper • 2608.02831 • Published 2 days ago • 5
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning Paper • 2508.08221 • Published Aug 11, 2025 • 50
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability Paper • 2405.14129 • Published May 23, 2024 • 14