Recent leaps in language model performance stem predominantly from verifiable agentic training data rather than architectural revisions alone. However, generating verifiable tasks for supervised fine-tuning and reinforcement learning still relies heavily on expensive human-in-the-loop engineering, hindering recursive self-improvement (RSI). Researchers from CUHK and collaborating organizations introduce AutoDataBench, the first benchmark evaluating an agent's capability to autonomously synthesize verified training tasks against industrial production acceptance standards. Unlike traditional evaluation schemes requiring expensive end-to-end retraining runs, AutoDataBench directly inspects generated tasks across validity, novelty, difficulty, and behavioral coverage. Evaluated on three executable benchmarks, existing frontier agents score below 20/100 within a 45-minute budget; however, quadrupling the test-time budget significantly boosts success rates while maintaining unit costs, proving that autonomous synthetic data pipelines are viable when scaled with compute.
Key Takeaways
- ✓Industrial Acceptance Criteria for Synthetic Tasks: Replaces expensive end-to-end retraining evaluations with a sample-by-sample automated harness assessing validity, novelty, difficulty, and coverage directly.
- ✓Frontier Agent Synthesis Baseline: Across three executable benchmarks, current top agents achieve under 20/100 at standard 45-minute budgets, demonstrating that synthesizing verifiable tasks is significantly harder than solving them.
- ✓Validating Recursive Self-Improvement Scaling: Quadrupling the agent's generation and self-testing time substantially elevates valid task yields while marginal unit costs remain flat, confirming the viability of compute-driven synthetic data engines.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.