Vals AI released MysteryMechanism, a benchmark testing whether agents can rediscover sealed scientific mathematical mechanisms via bounded experiments. GPT-6 Astra led at 53.2% accuracy versus Claude Fable 5.1 at 47.8% and Claude Opus 5 at 37.4%; Astra also cost about $1.77 per test, roughly one-third of second-place Fable 5.1 at $5.63.
Key Takeaways
- βMysteryMechanism tests rediscovering sealed scientific mechanisms via bounded experiments
- βGPT-6 Astra leads at 53.2% vs Claude Fable 5.1 47.8% and Claude Opus 5 37.4%
- βAstra costs about $1.77/test, roughly one-third of second-place Fable 5.1 at $5.63
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.