Covering Human Action Space for Computer Use: Data Synthesis and Benchmark Paper • 2605.12501 • Published May 12 • 17
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering Paper • 2603.20193 • Published Mar 20 • 1
LLMSurgeon: Diagnosing Data Mixture of Large Language Models Paper • 2605.30348 • Published May 28 • 1
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense Paper • 2602.09012 • Published Feb 9
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems Paper • 2604.14228 • Published Apr 14 • 25
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 64
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 64
LLMSurgeon: Diagnosing Data Mixture of Large Language Models Paper • 2605.30348 • Published May 28 • 1
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense Paper • 2602.09012 • Published Feb 9
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Paper • 2505.24878 • Published May 30, 2025 • 23
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published Aug 13 • 64
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting Paper • 2602.17645 • Published Feb 19
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models Paper • 2410.13859 • Published Oct 17, 2024 • 9