K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts Paper • 2606.02404 • Published Jun 1 • 59
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback Paper • 2605.17448 • Published May 17 • 19
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs Paper • 2605.09063 • Published May 9 • 82
OfficerChul/Qwen3-VL-30B-A3B-Instruct-Android-Control-84k Image-Text-to-Text • 31B • Updated Oct 26, 2025 • 16 • 1
OfficerChul/Qwen3-VL-30B-A3B-Instruct-Android-Control-84k Image-Text-to-Text • 31B • Updated Oct 26, 2025 • 16 • 1
OfficerChul/gemma-3n-E2B-it-Android-Control-84k Image-Text-to-Text • 5B • Updated Oct 16, 2025 • 17 • 2
OfficerChul/gemma-3n-E2B-it-Android-Control-84k Image-Text-to-Text • 5B • Updated Oct 16, 2025 • 17 • 2
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI Paper • 2510.05684 • Published Oct 7, 2025 • 147