Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 139 SWE-bench-Live/OS-bench Viewer • Updated 10 days ago • 140 • 299
SWE-bench-Live The datasets for benchmarking and training of LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated Sep 18, 2025 • 3.69k • 9.8k • 7 SWE-bench-Live/MultiLang Viewer • Updated May 16 • 743 • 2.19k SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 139
Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 139 SWE-bench-Live/OS-bench Viewer • Updated 10 days ago • 140 • 299
SWE-bench-Live The datasets for benchmarking and training of LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated Sep 18, 2025 • 3.69k • 9.8k • 7 SWE-bench-Live/MultiLang Viewer • Updated May 16 • 743 • 2.19k SWE-bench-Live/Windows Viewer • Updated Jun 28 • 61 • 139