MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 20 days ago • 77
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 17 days ago • 142
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Paper • 2604.11784 • Published Apr 13 • 143