HarnessEvolve decouples execution from evolution and aligns failed runs against reference trajectories (produced with ground-truth answers) to diagnose first-error steps, cluster recurring failure modes, then edit the harness—prompts, skills, tools, and execution logic—behind quality and performance gates. On CloudCoreNetwork-QA with Qwen3.6-27B, the full system hit 86.9% (57.8% without reference trajectories) and beat the strongest baseline by 21.6 points.
Key Takeaways
- ✓Method: align failures to reference trajectories to fix long-horizon credit assignment.
- ✓Edit surface: prompts, skills, tools, scripts, and execution logic behind leak/bloat and regression gates.
- ✓Result: 86.9% on CloudCoreNetwork-QA with Qwen3.6-27B (57.8% without refs); +21.6pp vs strongest baseline.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.