The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems.
SWE-bench-Live
community
AI & ML interests
None defined yet.
Recent Activity
Organization Card
SWE-bench-Live
Here we host SWE-bench-Live dataset, with continuous monthly updates!
models 0
None public yet
datasets 5
SWE-bench-Live/Windows
Viewer • Updated • 66 • 1.57k
SWE-bench-Live/MultiLang
Viewer • Updated • 1.08k • 21.9k
SWE-bench-Live/SWE-bench-Live
Viewer • Updated • 3.69k • 75.5k • 8
SWE-bench-Live/repo-build-test-benchmark
Viewer • Updated • 119 • 790
SWE-bench-Live/OS-bench
Viewer • Updated • 140 • 1.17k