On Oct 9 Pine AI released Pine Computer in private beta: cloud virtual computers plus a harness and runtime built for AI, driven by an API across browser, files and shell. Instead of repeated screenshots, the OS, browser and apps push change notifications to the model ("epoll for AI perception"). On SaaS-Bench v1.1 (106 tasks, 23 apps) GPT-5.6 Luna on Pine scored 78.3% on checkpoints vs 74.3% for Opus 5 + Claude Code, though its fully-resolved rate (27.4%) trails the latter (31.1%).

Key Takeaways

  • ✓SaaS-Bench v1.1 checkpoint score: Pine + GPT-5.6 Luna 78.3% vs Opus 5 + Claude Code 74.3% vs GPT-5.6 Sol + Codex 71.1%
  • ✓Fully resolved rate trails: 27.4% vs 31.1% (Claude Code) / 29.2% (Codex)
  • ✓Model token cost per task: ~$1.02 vs $26.50 (Opus 5 + Claude Code) / $20.50 (GPT-5.6 Sol + Codex), excluding infrastructure
  • ✓New computer boots in seconds, heavy state restores in ~15 s, paused ones resume in <1 s
  • ✓Launch post: ~99.8K impressions, 250+ likes, 120 reposts; private beta waitlist only
Pine Computer launches a cloud computer built for AI agents: event-driven perception instead of screenshot polling; GPT-5.6 Luna hits 78.3% SaaS-Bench v1.1 checkpoint score vs 74.3% for Opus 5 + Claude Code at ~$1.02 model cost per task
🖼️Official Media
Click to view high-res
🧭

Heavy Claude Code use: compare subscription limits and API bills

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Pine AI released Pine Computer, a private-beta cloud computer designed for AI rather than humans. It bundles virtual computers, a harness, a runtime wired into the OS and browser, and local/remote tools behind an SDK and API: an application submits a task, and the system works across browser, files and shell, returning progress and results. The core change is perception: instead of re-reading screenshots or accessibility trees after every step, the OS, browser and apps push change notifications to the model (the team calls it epoll for AI). Each computer is isolated with scoped access, saved state is encrypted with the customer's key, and users can take over the streamed screen for logins.

On SaaS-Bench v1.1 (106 tasks, 23 apps), GPT-5.6 Luna on Pine posted a 78.3% checkpoint score vs 74.3% for Opus 5 + Claude Code and 71.1% for GPT-5.6 Sol + Codex, but its fully-resolved rate (27.4%) was below both (31.1% and 29.2%). Model token cost was about $1.02 per task vs $26.50 and $20.50, excluding infrastructure. Pine notes these are whole-system comparisons with different budgets, and its 2–5x speed claim comes from internal preliminary tests.

Access is via waitlist; it currently runs Pine's own model or GPT-5.6 Luna, bring-your-own-model is planned, and it runs only in Pine's cloud. New computers start in seconds, heavy state restores in about 15 seconds, and paused computers resume in under a second.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.