OpenAI's first vertical-SaaS computer-use collaboration: with Ironclad it defined 11 legal, procurement and commercial tasks, let models practice in hosted Ironclad environments, and used RL on synthetic tasks. GPT-6 Astra scores 55.0% mean rubric vs 41.6% for GPT-5.6 Sol, with estimated time per attempt down from 37.0 to 19.2 minutes; an internal model reaches 63.7%. OpenAI is inviting more software companies to partner.

Key Takeaways

  • ✓11 real contracting tasks, each graded on 8–50 criteria; ~30–40 min per task for an experienced human user
  • ✓GPT-6 Astra (Max) 55.0% mean rubric vs GPT-5.6 Sol (High) 41.6%, ~32% relative gain
  • ✓Estimated time per attempt 37.0 → 19.2 minutes (~48% lower) — simulated, not measured customer savings
  • ✓Demo task: Astra ~94% of criteria in ~20 min vs Sol ~85% in ~32 min; an internal model hits 63.7%
  • ✓Synthetic training tasks built from public SEC EDGAR contracts; no OpenAI or Ironclad customer data used
OpenAI and Ironclad train GPT-6 Astra on real contracting software: 55.0% vs GPT-5.6 Sol's 41.6% on 11 workflows, ~48% less time per attempt
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

OpenAI's first published vertical-SaaS computer-use collaboration. With Ironclad, it picked 11 legal, commercial and procurement tasks (NDA setup, threshold-based finance approvals, jurisdiction-aware reusable clauses), graded each on 8–50 criteria, and let models practice in hosted Ironclad environments using RL on synthetic tasks built from filtered public SEC EDGAR contracts. GPT-6 Astra (Max reasoning) averaged 55.0% vs 41.6% for GPT-5.6 Sol (High), with estimated time per attempt falling from 37.0 to 19.2 minutes; an internal model reached 63.7%. Caveats: only 11 research tasks, and times are simulated from model speeds rather than measured customer savings. OpenAI is now inviting a small number of software companies to bring hard, verifiable workflows, secure test environments and research-safe data.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.