Artificial Analysis independently benchmarked Devin Fusion—the first multi-model coding agent on their Coding Agent Index. Claude Fable 5.1 (xhigh)+SWE-2 scores 62; GPT-6 Astra (xhigh)+SWE-2 scores 59 with ~43% lower cost and ~31% faster completion while retaining frontier-level performance.
Key Takeaways
- ✓First multi-model coding agent on Artificial Analysis Coding Agent Index.
- ✓Fable 5.1+SWE-2 scores 62; Astra+SWE-2 scores 59.
- ✓Astra config ~43% cheaper and ~31% faster task completion.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.