Arena.ai says Anthropic Claude Fable 5.1 (Max) is #1 on Agent Arena across 6.7k+ real-world long-horizon agentic sessions with +15.8% net improvement, redrawing the price-performance frontier at a $4.14 median cost/task. Implicit praise-vs-complaint leads by +42.5%, confirmed success +22.4%, bash recovery +13.1%, and no tool hallucinations; GPT-6 Astra traces are still coming in.

Key Takeaways

  • Claude Fable 5.1 (Max) tops Agent Arena with +15.8% net over 6.7k+ real agentic sessions.
  • Median cost $4.14/task — both the strongest and costliest model on the board, reshaping the Pareto frontier.
  • Praise-vs-complaint +42.5%, confirmed success +22.4%, bash recovery +13.1%, no tool hallucinations; Astra traces still arriving.
ADSponsored