Arena.ai says Anthropic Claude Fable 5.1 (Max) is #1 on Agent Arena across 6.7k+ real-world long-horizon agentic sessions with +15.8% net improvement, redrawing the price-performance frontier at a $4.14 median cost/task. Implicit praise-vs-complaint leads by +42.5%, confirmed success +22.4%, bash recovery +13.1%, and no tool hallucinations; GPT-6 Astra traces are still coming in.
Key Takeaways
- ✓Claude Fable 5.1 (Max) tops Agent Arena with +15.8% net over 6.7k+ real agentic sessions.
- ✓Median cost $4.14/task — both the strongest and costliest model on the board, reshaping the Pareto frontier.
- ✓Praise-vs-complaint +42.5%, confirmed success +22.4%, bash recovery +13.1%, no tool hallucinations; Astra traces still arriving.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.