Vals AI released MysteryMechanism, a benchmark testing whether agents can rediscover sealed scientific mathematical mechanisms via bounded experiments. GPT-6 Astra led at 53.2% accuracy versus Claude Fable 5.1 at 47.8% and Claude Opus 5 at 37.4%; Astra also cost about $1.77 per test, roughly one-third of second-place Fable 5.1 at $5.63.

Key Takeaways

  • βœ“MysteryMechanism tests rediscovering sealed scientific mechanisms via bounded experiments
  • βœ“GPT-6 Astra leads at 53.2% vs Claude Fable 5.1 47.8% and Claude Opus 5 37.4%
  • βœ“Astra costs about $1.77/test, roughly one-third of second-place Fable 5.1 at $5.63
ADSponsored