The SWE-bench benchmark team launched SWE-bench Pro, expanding from Python to Rust, Go, C++, and TypeScript. Featuring 1,200 curated real-world issues from high-traffic open-source repositories, it provides a unified industry standard for multi-language autonomous agent software engineering evaluation.
Key Takeaways
- βIncludes 1,200 real-world repository repair tasks across Rust, Go, C++, and TypeScript.
- βApplies strict compiler verification and integration test matrices to eliminate false passes.
- βFirst leaderboard standings show Claude 3.7 Sonnet and DeepSeek-R1 leading multi-language solve rates.