CoArena (@coastyai, YC S26) released CUA KnowledgeBench (Real Company Work): a 12-hour horizon with 10 private real-world tasks × 5 seeds across 100+ apps, scored pass/fail on the work product. After spending over $50k on frontier models, Claude Fable 5.1 solved 12 of 50 twelve-hour runs (24%) while GPT-6 Astra barely cracked the top 3. Next they plan $200k+ compute for 1-, 3-, and 7-day horizons.

Key Takeaways

  • New CUA benchmark targets 12-hour real company knowledge work.
  • Pass/fail on work product across 100+ apps and multiple seeds.
  • Fable 5.1 ~24% pass; Astra barely top-3; longer horizons coming.
ADSponsored