Qwen Team's Terminal-Universe reconstructs reusable executable workspaces from terminal/code-agent trajectories via deterministic replay plus an agentic completer, then synthesizes new tasks along breadth (cross-workspace) and depth (multi-round user sessions) with executable verifiers. On public trajectories it yields ~37.3k task-sufficient environments and ~32k SFT demos; SFT of Qwen3.5-27B improves Terminal-Bench 2.1 by 11.9 points and EvoCode-Bench v2 MT@4 by 13.8 points—beating imitation of raw trajectories.
Key Takeaways
- ✓Rebuild: replay file ops then agentically complete missing deps into reusable verifiable workspaces.
- ✓Scale: ~37.3k task-sufficient envs; ~32k Single-WS / Cross-WS / Multi-Round SFT demos.
- ✓Gains: Qwen3.5-27B +11.9 on Terminal-Bench 2.1 and +13.8 on EvoCode-Bench v2 MT@4.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.