Epoch AI evaluated GPT-5.6 Sol on Slay the Spire via computer-use plugins, finding that while the model reliably commands the UI and clears low ascensions (A4-A9), long-term strategic decision-making remains a key hurdle.
Key Takeaways
- ✓Computer-use vision execution is reliable, successfully identifying complex UI elements and state transitions;
- ✓Strategic planning latency is high, requiring 5-8 hours of thinking per run with occasional tactical misplays;
- ✓Epoch initiated follow-up benchmarks on Slay the Spire 2 with live streamed autonomous agent runs on Twitch.