GitHub launched Project HydraFusion as a research preview: it orchestrates models and workflows per coding task. In controlled offline evals vs Claude Opus 5 on Terminal-Bench 2.1 it reported higher verified task quality and lower estimated cost. Select it like any other model in GitHub Copilot, or try via /experimental in Copilot CLI.
Key Takeaways
- ✓Bench: ~+4.9 percentage points verified task quality vs Claude Opus 5 on Terminal-Bench 2.1.
- ✓Cost: ~67% lower estimated workflow cost in controlled offline evaluations.
- ✓Access: research preview in GitHub Copilot; try via /experimental in Copilot CLI.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.