RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 13 days ago • 141
JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation Paper • 2606.03168 • Published Jun 2 • 47
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model Paper • 2511.13647 • Published Nov 17, 2025 • 72