Recent attempts to scale tool-use post-training for autonomous agents have focused primarily on synthesizing standalone execution environments. However, environments constitute merely one component of a holistic agentic interaction system comprising environments, tasks, harnesses, and evaluators; scaling environments in isolation yields inconsistent training signals. Researchers from Fudan University introduce WEFT (arXiv:2609.36887), a whole-system evolution framework for tool-use post-training. WEFT coordinates environment breadth, task complexity, and interaction diversity, applying execution-driven self-evolution to attribute failures and iteratively revise flawed components. To ensure optimization stability and concurrency reliability, WEFT introduces prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP—a Model Context Protocol infrastructure providing isolated, recoverable state management across concurrent rollouts. WEFT-8B and WEFT-14B outperform matched environment-scaling baselines across BFCL V4, tau^2-Bench, and Claw-Eval, with WEFT-14B beating Agent-World-14B by 6.41, 2.23, and 12.27 percentage points, while WEFT-35B-A3B excels on long-horizon benchmarks like Toolathlon-Verified.

Key Takeaways

  • ✓Pioneers WEFT, co-evolving environments, tasks, harnesses, and evaluators to overcome the limits of isolated environment synthesis
  • ✓Integrates prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP isolation for resilient tool-use post-training
  • ✓WEFT-14B surges past Agent-World-14B by 6.41 pp on BFCL V4, 2.23 pp on tau^2-Bench, and 12.27 pp on Claw-Eval
WEFT: Whole-System Evolution for Tool-Use Post-Training Scales Agentic Interaction with MegaMCP State Isolation
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Scaling tool-use capabilities through post-training is foundational for modern autonomous coding and enterprise agents. However, recent paradigms primarily emphasize synthesizing executable sandbox environments (such as Agent-World). In practice, an environment is only one constituent of an agentic interaction system that equally depends on nuanced task definitions, execution harnesses, and robust evaluators. Scaling environments in isolation produces severe cross-component incoherence, generating noisy reward signals that destabilize policy optimization and trigger overfitting to synthetic quirks.

架构亮点与底层机制

Researchers from Fudan University present WEFT (Whole-system Evolution For Tool-use Post-training, arXiv:2609.36887):

  1. Whole-System Co-Construction: Jointly scales the entire agentic topology across environmental diversity (spanning diverse enterprise APIs), task complexity, and multi-turn interaction modalities.
  2. Execution-Driven Self-Evolution: Inspects full execution traces and runtime state invariants to pinpoint root causes behind failures across environments, harnesses, or rubrics, applying targeted updates verified by follow-up rollouts.
  3. Stable Post-Training Optimizations: Introduces prefix-preserving sampling to safeguard confirmed progress during RL exploration, coupled with atomic-turn credit assignment to concentrate policy gradients on critical tool-invocation decisions.
  4. MegaMCP Infrastructure: Integrates with the Model Context Protocol (MCP) standard to orchestrate isolated, snapshot-recoverable state management across massive concurrent agent rollouts over shared backend services.

权威 Benchmark 与实测跑分对比

Empirically evaluated against state-of-the-art environment-scaling baselines:

  1. Dominates Across BFCL V4, tau^2-Bench, and Claw-Eval: Both WEFT-8B and WEFT-14B establish new performance highs against identically sized foundation models.
  2. Up to +12.27 Percentage Point Gains Over Agent-World-14B: WEFT-14B outperforms Agent-World-14B by 6.41 percentage points on BFCL V4, 2.23 points on tau^2-Bench, and an extraordinary 12.27 percentage points on Claw-Eval.
  3. Robust Long-Horizon Scaling: The scaled WEFT-35B-A3B variant maintains superior accuracy on arduous long-horizon task benchmarks, including Toolathlon-Verified and AutomationBench.

开发者实战落地与开箱指南

WEFT provides a definitive blueprint for machine learning engineers building tool-using agents. By replacing isolated environment generation with whole-system evolution and deploying MegaMCP for resilient state management, teams can construct stable, sample-efficient reinforcement learning pipelines that transfer reliably to production APIs and complex multi-agent workflows.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.