Modern software engineering transcends isolated code edits: engineers run applications, interact with GUI interfaces, and visually inspect rendering feedback to diagnose errors and verify changes. CUA-SWE bridges coding agents and computer-use agents by introducing the first unified benchmark and environment for visual software engineering. Spanning four engineering domains, CUA-SWE challenges agents to interleave code/config modifications, terminal execution, GUI interactions, and screenshot inspections, with deterministic test suites evaluating behavioral correctness.

Key Takeaways

  • ✓First multimodal visual software engineering benchmark coupling computer-use capabilities with code repositories
  • ✓Text-only agents stall below 12% on visual defect repairs, while closed-loop visual feedback elevates success to 43.8%
  • ✓Identifies attention contention bottlenecks between high-resolution GUI screenshots and repository-scale source code
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 Traditional coding benchmarks like SWE-bench assume specifications and verification are purely text-based. In industry, visual defects (layout glitches, interactive race conditions, canvas rendering) cannot be diagnosed or resolved without launching applications, manipulating interfaces, and visually inspecting state changes. Coding agents and computer-use agents (CUA) have remained siloed, leaving visual software engineering unaddressed. ### 架构亮点与底层机制 CUA-SWE establishes a unified testbed combining code manipulation with interactive visual environments: 1. Integrated Multi-Modal Sandbox: Equips agents with terminal access, code editing tools, and direct graphical interaction (mouse/keyboard events, screenshot feeds). 2. Visually-Gated Specifications: Tasks purposefully place essential requirements exclusively within active UI interfaces, compelling agents to interact and inspect runtime behavior. 3. Visual-Code Causal Mapping: Assesses whether agents can trace visual UI artifacts back to underlying component lifecycles, CSS properties, or state machines. 4. Deterministic Dual-Layer Verification: Couples behavioral GUI interaction suites with standard unit test harnesses to guarantee functional compliance. ### 权威 Benchmark 与实测跑分对比 Evaluated across Web, Desktop GUI, Interactive Dashboards, and Game Tooling: 1. Text-Only Agents Suffer Severe Degradation: Deprived of GUI feedback, frontier models solve less than 12% of visually-dependent bug repairs. 2. Closed-Loop Visual Feedback Elevates Success to 43.8%: Interleaving visual inspection with incremental code edits propels success rates to 43.8% for leading frontier models. 3. Attention Bottlenecks Identified: Dense visual tokens across extended interaction horizons create context contention against repository-scale code navigation. ### 开发者实战落地与开箱指南 CUA-SWE is open-sourced on GitHub with containerized environments and automated runners. Teams building next-generation multimodal IDEs can leverage CUA-SWE to train and validate agents that possess both code-generation and visual-verification capabilities.