OpenAI says GPT-6 Astra is state-of-the-art on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro—benchmarks for cross-profession computer workflows and screen grounding—adding computer-use evidence beyond the already announced TerminalBench / ARC-AGI / FrontierMath results.
Key Takeaways
- ✓Claims SOTA on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro for computer workflows/screen ops.
- ✓Framed as cross-profession computer use, not just repo-fix coding.
- ✓Complements already-announced TerminalBench / ARC-AGI / FrontierMath agentic evidence.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.