Cross-platform-bench Collection The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems. • 2 items • Updated Jun 28
SWE-bench-Live Collection The datasets for benchmarking and training of LLM coding agents. • 3 items • Updated Jan 22 • 1