Specific Labs (@janaksunil) launched Real-SWE, evaluating coding agents on private, out-of-distribution production codebases from real companies (200K+ user apps, fintech statement pipelines, enterprise sales tools, and more). On the public board (8-run averages) Claude Fable 5.1 + Claude Code leads at 38.8%, GPT-6 Astra + Codex CLI at 33.8%, and Gemini 3.8 Flash + Gemini CLI at 31.2% — none clear 40%. Site: realswe.withspecific.com.
Key Takeaways
- ✓Tasks come from real private company codebases — billing, permissions, customer data — not public contest repos.
- ✓8-run averages: Fable 5.1 + Claude Code 38.8%, Astra + Codex CLI 33.8%, Gemini 3.8 Flash + Gemini CLI 31.2%.
- ✓Built for the systems companies would actually deploy agents into; board at realswe.withspecific.com.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.