Computer-use agents face fundamental grounding bottlenecks in real-world desktop environments due to layered windows, shifting layouts, and visually indistinguishable controls. Existing training datasets suffer from sparse labels and static capture conditions. Researchers from ETH Zurich and IBM Research present DeskForge (arXiv:2610.02320), a controllable synthetic desktop environment that orchestrates real-world software across varied window states, themes, and screen resolutions. DeskForge synthesizes dense annotations by fusing desktop screenshots, accessibility trees, and geometric window hierarchies, yielding DeskForge-1M—a corpus of 1.2M observations and 159.7M annotated UI element instances. Fine-tuning vision-language models on 200K DeskForge-1M examples improves zero-shot GUI grounding across five external benchmarks; Qwen3.5-4B gains 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. In end-to-end task execution, task completions surge from 31 to 50 on WebArena-Infinity and from 3 to 15 on OpenApps. The dataset, models, and code are open-sourced at saidgurbuz.github.io/deskforge.

Key Takeaways

  • ✓ETH Zurich and IBM introduce DeskForge, a controllable environment synthesizing DeskForge-1M with 1.2M observations and 159.7M UI elements
  • ✓Fine-tuning Qwen3.5-4B boosts GUI grounding by 11.51 points on ScreenSpot-Pro and 10.11 points on OSWorld-G across five external benchmarks
  • ✓Multiplies long-horizon task completion: solves 50/119 tasks on WebArena-Infinity (up from 31) and 5x on OpenApps (from 3 to 15), fully open-sourced
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

Operating-system-level computer-use agents must reliably localize targets across messy desktop ecosystems where overlapping windows, non-standard GUI controls, dynamic resolutions, and visual clutter create acute grounding ambiguity. Static web datasets or mobile traces do not capture arbitrary window occlusion, desktop context menus, or floating dialog states, causing agents to target background windows or miscalculate interactive coordinates.

Architecture and How It Works

DeskForge (arXiv:2610.02320) addresses this bottleneck by orchestrating controllable desktop synthetic pipelines:

  1. Programmatic Desktop Sandboxing: Operates within real OS environments, algorithmically generating variations in application state, layout hierarchies, color schemes, and display resolutions.
  2. Tri-Source Dense Fusion: Simultaneously synchronizes pixel screenshots, OS-level accessibility trees, and window manager geometry to derive bounding boxes and interaction metadata without manual annotation.
  3. DeskForge-1M Corpus: Synthesizes 1.2M desktop observations featuring 159.7M annotated UI element instances alongside paired pre- and post-action delta traces.

Benchmarks and Measured Results

Vision-language models fine-tuned on a 200K subset of DeskForge-1M demonstrate dramatic capabilities:

  1. Grounding Across Five Benchmarks: Qwen3.5-4B improves by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G, establishing state-of-the-art open-weight performance.
  2. End-to-End Task Success: In long-horizon execution under frozen planners, fine-tuned models complete 50 out of 119 WebArena-Infinity tasks (vs. 31 originally) and 15 out of 100 OpenApps workflows (a 5x improvement over 3).
  3. Zero-Shot Application Robustness: Transfers robustly across unobserved proprietary enterprise software and customized UI elements.

Getting Started for Developers

The DeskForge generation engine, dataset, and fine-tuned checkpoints are public at saidgurbuz.github.io/deskforge. Teams building RPA copilots or OS agents can deploy the fine-tuned vision models directly for click grounding or deploy the sandbox to synthesize synthetic training pairs for proprietary desktop applications.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.