Artificial Analysis benchmarks MBZUAI’s new open-weights MoE K2 Horizon 375B A23B (375B total / 23B active, 512K context) at 47 on the Intelligence Index—up sharply from dense K2 Think V2 (17). It leads nearby MiniMax-M3 on agentic/knowledge-work (GDPval-AA, τ³-Banking) but trails on hard reasoning (GPQA Diamond, HLE); ~26% hallucination rate is driven by abstention more than knowledge.
Key Takeaways
- ✓Intelligence Index 47 vs K2 Think V2 at 17; 70B dense → 375B/23B MoE with 512K context.
- ✓Agentic: GDPval-AA Elo 1430 vs MiniMax-M3 1380; τ³-Banking 34.2% vs 15.3%.
- ✓Weaker on hard knowledge/reasoning (GPQA/HLE); AA-Omniscience attempts ~40% of questions. Pricing and license TBD.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.