Endowing LLM agents with test-time self-improvement—the ability to iteratively refine solutions—is a vital milestone for autonomous engineering. Researchers from Renmin University and collaborators present AREX-2, an open-source framework leveraging long-horizon reflective trajectories. Decoupling self-improvement into reflection (generating superior alternatives) and long-horizon execution (sustaining iteration over multiple turns), AREX-2 trains on verified machine learning and algorithmic programming trajectories. Built on Qwen3.8-27B, AREX-2 achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, transferring effectively to deep research with 92.2 on GAIA and 93.8 on DeepSearchQA.

Key Takeaways

  • ✓Decouples test-time self-improvement into domain-agnostic reflection and long-horizon execution
  • ✓Achieves 81.8 on MLE-bench Lite and 70.7 on Frontier-CS with Qwen3.8-27B, transferring to 92.2 on GAIA
  • ✓Breaks multi-turn degradation bottlenecks, showing monotonic score improvements as iteration budgets increase
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 Most autonomous agents follow brittle, single-pass generation patterns. Complex engineering challenges require continuous trial-and-error, reflective self-critique, and iterative refinement. Previous efforts in self-correction struggled beyond 1-2 turns, often succumbing to context degradation, goal drift, and degenerative loops. ### 架构亮点与底层机制 AREX-2 formalizes test-time self-improvement into two domain-agnostic pillars: reflection (generating hypotheses superior to the incumbent solution) and long-horizon execution (maintaining stability across tens of iterative cycles). Recognizing that machine learning and algorithmic programming provide deterministic execution verifiers (loss curves, metrics, unit tests), the authors synthesized thousands of long-horizon reflective trajectories to align Qwen3.8-27B. ### 权威 Benchmark 与实测跑分对比 Evaluated on demanding programming and open-domain agent benchmarks: 1. 81.8 on MLE-bench Lite & 70.7 on Frontier-CS: AREX-2 sets impressive records on real-world Kaggle-style ML pipelines and complex algorithmic challenges. 2. Seamless Transfer to Deep Research: Long-horizon reflective capability transfers zero-shot to deep web research, scoring 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA. 3. Monotonic Scaling with Reflection Rounds: Performance scales monotonically as test-time reflection budgets expand, eliminating the regression collapse common in multi-turn prompting. ### 开发者实战落地与开箱指南 AREX-2 is open-sourced on GitHub with ready-to-use iterative execution runners. Engineers building automated CI/CD refactoring or autonomous Kaggle pipelines can deploy AREX-2 with defined unit-test verifiers to unlock long-horizon, autonomous code refinement.