Developer 3s Key Decision Metrics
An unsuccessful agent rollout contains rich diagnostic signals beyond sparse negative rewards. UIUC researchers introduced the Agent Error Dataset (AED), the largest scale failure dataset comprising 50,228 error-diagnosis pairs across 9,961 source tasks, 33 environments, 19 harness families, and 23 policy models. Powered by a 5-stage Agentic Error-to-Training (AET) pipeline, first-proposal repairs lift verifier pass rates from 18.4% to 51.1% (+32.7 percentage points). Furthermore, post-training Qwen3-8B on AED boosts exact-step failure diagnosis agreement from 47.2% to 63.6%.
Key Takeaways
- ✓Features 50,228 curated error-diagnosis pairs across 33 environments, 19 harnesses, and 23 policy models
- ✓Lifts replay verifier pass rates from 18.4% to 51.1% (+32.7 percentage points) with first-proposal corrections
- ✓Boosts Qwen3-8B exact-step failure diagnosis agreement from 47.2% to 63.6%, demonstrating superiority over success-only training
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.