Existing Computer-Use Agents (CUAs) face a dilemma: relying strictly on graphical user interface (GUI) interactions is slow and error-prone, while building per-application API tools demands unsustainable engineering and lacks generalizability. Researchers from Zhejiang University introduce HybridCUA, combining the UI versatility of GUIs with the execution speed and determinism of the Command Line Interface (CLI). Supported by the newly constructed HybridCUA-8K dataset (5K hybrid trajectories and 3K RLVR tasks), HybridCUA employs a two-stage training paradigm: supervised fine-tuning across interleaved GUI-CLI modalities, followed by CLI-aware reinforcement learning that rewards reliable, selective shell execution. On the OSWorld benchmark, HybridCUA-9B achieves 53.6% accuracy, improving over the base model by 14.8 percentage points, while delivering a 4.0-point gain on WindowsAgentArena.
Key Takeaways
- ✓Unified GUI and CLI Paradigm: Bridges the gap between fragile vision-only screen clicking and application-specific APIs, combining the speed of shell commands with the visual ubiquity of desktop GUIs.
- ✓HybridCUA-8K Trajectory Dataset: Curates 5,000 interleaved GUI-CLI trajectories and 3,000 verified RLVR tasks with ground-truth environment state feedback for agent post-training.
- ✓53.6% on OSWorld with Cross-OS Generalization: HybridCUA-9B hits 53.6% accuracy on OSWorld (a 14.8 percentage point gain over base) and boosts WindowsAgentArena by 4.0 points across distinct OS environments.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.