Artificial Analysis (@ArtificialAnlys) shipped Intelligence Index v4.3: Terminal-Bench upgraded 2.1→4.0 (now mini-SWE-agent harness) and τ³-Banking replaced by AutomationBench-AA (657 private Zapier business workflows). Claude Fable 5.1 and GPT-6 Astra tie at 53; OpenAI holds most of the intelligence-vs-cost frontier across Astra reasoning efforts.
Key Takeaways
- ✓Benchmarks: Index v4.3 upgrades Terminal-Bench to 4.0 and adds AutomationBench-AA (private set).
- ✓Weighting: private-task/answer evals rise from 40% to 45%; category weights unchanged.
- ✓Results: Claude Fable 5.1 and GPT-6 Astra tie at 53; open-weights led by GLM-5.3 / Kimi K3 at 44.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.