Artificial Analysis shipped Intelligence Index v4.2: adds private agentic knowledge-work eval AA-Briefcase and long-document GDP.pdf, doubles held-out weighting, upgrades grading, and drops saturated GPQA Diamond. Claude Fable 5.1 leads; GPT-6 Astra is second with ~4pt gain over GPT-5.6 Sol and strong output-token Pareto efficiency. A major interim benchmark refresh for model selection and cost/quality tradeoffs.
Key Takeaways
- βBench: v4.2 adds AA-Briefcase and GDP.pdf; held-out weight ~40%.
- βLeaderboard: Claude Fable 5.1 first, GPT-6 Astra second (~+4 vs Sol).
- βEfficiency: Astra stands out on the output-token Pareto frontier.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.