Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated 12 days ago • 66 • 1.48k SWE-bench-Live/OS-bench Viewer • Updated Jul 21 • 140 • 1.25k
SWE-bench-Live The datasets for benchmarking and training LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated 27 days ago • 3.69k • 78.1k • 9 SWE-bench-Live/MultiLang Viewer • Updated 12 days ago • 1.08k • 22.1k SWE-bench-Live/Windows Viewer • Updated 12 days ago • 66 • 1.48k
Cross-platform-bench The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. SWE-bench-Live/Windows Viewer • Updated 12 days ago • 66 • 1.48k SWE-bench-Live/OS-bench Viewer • Updated Jul 21 • 140 • 1.25k
SWE-bench-Live The datasets for benchmarking and training LLM coding agents. SWE-bench-Live/SWE-bench-Live Viewer • Updated 27 days ago • 3.69k • 78.1k • 9 SWE-bench-Live/MultiLang Viewer • Updated 12 days ago • 1.08k • 22.1k SWE-bench-Live/Windows Viewer • Updated 12 days ago • 66 • 1.48k