To overcome the severe GPU starvation and explosive trajectory redundancy of extreme-long (xlong) horizon agents spanning hours, hundreds of environment steps, and nearly 1M tokens per rollout, Alibaba's Qwen team introduced QwenGyre. By dynamically reallocating GPUs between rollout and training without pausing live interactions and pruning non-linear trajectory branches, QwenGyre enables 700K-token RL on the 2.4T-parameter Qwen 3.8 flagship, driving a 6.0% absolute gain on NL2RepoBench in 48 steps while delivering a 1.85x speedup.

Key Takeaways

  • ✓Breakthrough Infrastructure for xLong-Horizon RL: First production-grade RL framework built to sustain single-rollout trajectories spanning up to 1M tokens and hundreds of interactive steps for software engineering agents.
  • ✓Elastic Live GPU Reallocation: Eradicates massive GPU idle bubbles caused by execution variance by dynamically migrating GPUs between rollout and backward training clusters without halting live environment containers.
  • ✓2.4T-Parameter Model Gains & 1.85x Speedup: Scaled to the 2.4-trillion-parameter Qwen 3.8 backbone with 700K tokens per rollout, driving an absolute 6.0% accuracy gain on NL2RepoBench (52.5% to 58.5%) in 48 steps with 1.85x higher training throughput.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points As autonomous coding agents transition from isolated single-bug patches to repository-scale software development, task horizons have expanded by orders of magnitude (xLong-Horizon). Real-world autonomous workflows routinely span hours, hundreds of tool interactions, and nearly 1M tokens per rollout trajectory. Existing on-policy RL architectures (Colocate or decoupled Async) collapse under this regime: execution latency variance produces massive GPU idle bubbles during rollout waits, while non-linear trial-and-error branching creates intractable trajectory redundancy that cripples backward training passes. ### 架构亮点与底层机制 / Architectural Highlights Alibaba's Qwen team developed QwenGyre, an end-to-end framework specialized for xlong-horizon online RL: 1. Elastic Non-Interruptive GPU Reallocation: Seamlessly shifts GPU compute between rollout workers and gradient training nodes in real time without pausing or resetting active long-running environment sandboxes; 2. Branching Trajectory Processor: Reconstructs execution trees, prunes redundant trial-and-error execution loops, and assigns granular partial-progress credit to uncompleted trajectories to constrain memory footprints; 3. Megatoken Sequence Parallelism: Integrates ultra-long context attention distributed primitives capable of processing 700K+ token backpropagation updates per step on dense and MoE architectures. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Scaled to Alibaba's flagship Qwen 3.8 (2.4T parameters) with rollouts averaging 700,000 tokens: - NL2RepoBench Breakthrough: Achieves a +6.0% absolute accuracy leap (52.5% to 58.5%) on repository synthesis in only 48 RL steps; - Massive Training Speedup: Outperforms traditional Colocate architectures by 1.85x and decoupled Async pipelines by 1.78x across heterogeneous training domains; - Compute Efficiency: Maintains sustained GPU cluster computational efficiency above 88% despite extreme task length disparities. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Technical Reference: Detailed algorithms are available in arXiv preprint 2609.33848; - Paradigm Shift: Establishes that trillion-parameter foundation models can undergo stable online RL across megatoken horizons; - Ecosystem Integration: Core scheduling and trajectory deduplication components are being prepared for open-source integration within the Qwen ecosystem.