Addressing the limitation where synthetic training environments only generate tasks within isolated individual domains, researchers introduced CompoWorld. By scaling task spaces across 448 typed services exposing 10,130 tools and generating cross-service workflows via dependency graphs, paired with Completion-Focused Rubric Rewards in RL, a Qwen3.6-35B-A3B agent trained on just 3K SFT trajectories and 1K RL tasks gains +9.17 points across eight benchmarks and surpasses Claude Opus 4.6 on AutomationBench.
- ✓Compositional Scaling Across 10,130 Tools: Breaks the single-environment toy benchmark barrier by deploying coding agents to implement 448 verified microservices exposing 10,130 tools under unified state schemas.
- ✓Dependency Graph Random-Walk Workflows: Chains disparate services through dependency topologies to generate verifiable agent workflows where context and execution state must navigate multi-service boundaries.
- ✓35B Backbone Surpasses Frontier Models: Lifts eight agent benchmark averages by +9.17 points, overtaking Claude Opus 4.6 on AutomationBench and setting a new efficiency record for open-weights agent models.
🧭Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
🔗
Project Links & Resources
Direct AccessDirect access to official project resources and documentation🔬
In-Depth Technical Analysis
核心背景与行业痛点 / Background & Pain Points Autonomous agents excel at isolated, single-domain tool calls but collapse in enterprise production workflows where decisions require orchestrating state and data flows across disparate third-party services. Existing synthetic agent environments generate tasks within isolated silos, while manually designing multi-service end-to-end integration datasets remains prohibitively expensive and unscalable. ### 架构亮点与底层机制 / Architectural Highlights Researchers from Nanjing University and collaborating labs formulated CompoWorld: 1. Compositional Service Scaling: Employs coding agents to instantiate 448 typed microservices exposing 10,130 discrete tools with uniform state interfaces, augmented by world-model fallbacks for non-deterministic web services; 2. Random-Walk Dependency Workflows: Chains disparate tools via directed dependency graphs, synthesizing verifiable long-horizon scenarios where prerequisite information is generated by upstream services and consumed downstream; 3. Completion-Focused Rubric Reward: Directs policy gradient optimization in RL by dynamically prioritizing failure-prone transition criteria within rollout groups to enforce true end-to-end completion. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Fine-tuning Qwen3.6-35B-A3B with just 3K SFT trajectories and 1K RL tasks yielded striking empirical milestones: - Cross-Benchmark Gains: Achieves a +9.17 point mean gain across eight standard tool-use benchmarks over the base checkpoint; - Surpassing Frontier Models: On AutomationBench, CompoWorld outperforms Claude Opus 4.6 and leads all open and specialized 35B-scale agent architectures; - Sample Efficiency: Validates that combinatorial graph-driven task generation unlocks generalization without massive pre-training data volumes. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Technical Reference: Detailed algorithms are presented in arXiv preprint 2609.33665; - Open Assets: Service schemas, dependency generators, and evaluation rubrics are accessible via Hugging Face Papers; - Enterprise Takeaway: To train reliable integration agents, engineering teams should model cross-service API contracts and dependency trees rather than fine-tuning on disconnected single-API prompt pairs.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.