What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents Paper • 2610.06406 • Published 5 days ago • 15
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Paper • 2502.09925 • Published Feb 14, 2025
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing Paper • 2602.14186 • Published Feb 15
What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents Paper • 2610.06406 • Published 5 days ago • 15
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization Paper • 2603.29664 • Published Mar 31 • 50
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Paper • 2505.02835 • Published May 5, 2025 • 28