Computer-Use Agents (CUAs) designed to operate graphical desktop environments through virtual keyboard and mouse interactions are traditionally evaluated using outcome-based functional verifiers that assess only the final state after hundreds of operational turns. This black-box approach obscures whether failures stem from syntax typographical errors, GUI bounding-box misalignments, or intermediate state drift, preventing actionable diagnostic insights. NVIDIA Research introduces OSWorld-Pro (arXiv:2609.24890), the first comprehensive process-based evaluation suite for computer-use agents. Spanning 300+ complex desktop workflows decomposed into 2,800+ sequentially dependent subgoals grounded in 67,000+ expert human annotations, OSWorld-Pro deploys calibrated LLM judges to track incremental subgoal completion. Evaluations reveal that OSWorld-Pro presents a formidable challenge to state-of-the-art models: frontier models like Claude Opus 5 achieve only 75.7% (compared to 83.4% on standard OSWorld), systematically isolating critical procedural failure patterns such as exploratory irrelevant actions and click grounding drifts.

Key Takeaways

  • ✓NVIDIA Research unveils OSWorld-Pro, decomposing 300+ desktop tasks into 2,800+ sequentially dependent subgoals
  • ✓Grounded in 67,000+ expert annotations, revealing Claude Opus 5 scores only 75.7% vs 83.4% on standard OSWorld
  • ✓Isolates subgoal-irrelevant actions and click coordinate misalignments, replacing black-box outcome verifiers with process diagnostics
OSWorld-Pro: NVIDIA Unveils Process-Based Benchmark for Computer-Use Agents Spanning 2,800+ Subgoals
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Computer-Use Agents (CUAs)—such as Anthropic Computer Use and desktop automation assistants—interact directly with graphical operating systems via screenshots, mouse movements, and keyboard keystrokes. However, contemporary evaluation frameworks (exemplified by OSWorld) rely almost exclusively on outcome-based verifiers that inspect only the final environment state after hundreds of operational turns. This black-box methodology completely obscures where, when, and why an agent failed. A failure caused by a keyboard typographical error demands a fundamentally different post-training mitigation than one caused by a 5-pixel mouse click bounding-box offset, severely impeding algorithmic diagnostics.

架构亮点与底层机制

NVIDIA Research introduces OSWorld-Pro (arXiv:2609.24890), the first process-based benchmark for desktop CUAs:

  1. Granular Subgoal Decomposition: Spans 300+ real-world operating system tasks decomposed into over 2,800 sequentially dependent subgoals across office software, web browsing, and system configurations.
  2. Dense Human Expert Annotation: Grounded in over 67,000 rigorous human interaction annotations tracking operational states, mouse trajectories, and application transitions.
  3. Human-Aligned Multimodal LLM-Judges: Evaluates progressive subgoal fulfillment after every transitional action, tracking incremental completion velocity across long-horizon trajectories.
  4. Mechanistic Diagnostic Taxonomy: Pinpoints primary procedural failure archetypes, specifically isolating subgoal-irrelevant exploratory actions, UI grounding misclicks, and state synchronization lags.

权威 Benchmark 与实测跑分对比

Benchmarked across premier foundation models and agent frameworks:

  1. Strict Reality Check for Frontier Models: Even top-performing systems experience significant score drops under process scrutiny; Claude Opus 5 achieves only 75.7% on OSWorld-Pro compared to 83.4% on standard terminal-state OSWorld.
  2. Revealing Subgoal Progress Curves: While open-weight agents frequently diverge within the first 30% of a task due to irrelevant clicks, proprietary frontier models sustain alignment past 75% before faltering at multi-window state synchronization.
  3. Click Grounding as the Primary Bottleneck: Identifies that over 40% of execution errors stem directly from pixel-level coordinate targeting errors on high-DPI displays rather than high-level logical reasoning deficits.

开发者实战落地与开箱指南

OSWorld-Pro evaluation harnesses and annotation assets are available on the OSWorld repository (github.com/xlang-ai/OSWorld). Teams building OS automation agents, enterprise RPA copilots, and multimodal GUI foundation models can integrate OSWorld-Pro into their continuous integration pipelines to diagnose spatial grounding and long-horizon step-by-step reliability.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.