EVO-WAM: Evolving World Action Models through Video-Action Verification Paper • 2609.38057 • Published 1 day ago • 7
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Paper • 2406.18629 • Published Jun 26, 2024 • 42