JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 8 days ago • 122
Running on Zero MCP Featured 3.37k Wan2.2 14B Fast 🎥 3.37k generate a video from an image with a text prompt
MMSkills: Towards Multimodal Skills for General Visual Agents Paper • 2605.13527 • Published May 14 • 122
Running on CPU Upgrade 604 Visualize Dataset (v2.0+ latest dataset format) 💻 604 Explore and visualize LeRobot datasets
Running on Zero MCP Featured 406 Qwen Edit Any Pose 🕺 406 Edit any pose with Qwen Edit 2511 Any Pose LoRA
MolmoAct2: Action Reasoning Models for Real-world Deployment Paper • 2605.02881 • Published May 4 • 356
GUI-G^2: Gaussian Reward Modeling for GUI Grounding Paper • 2507.15846 • Published Jul 21, 2025 • 135