Together ran about 900 DeepSWE rollouts of GLM-5.3 vs GLM-5.3 Flash: 69.0% pass@1 at ~$3.99 per task versus 63.4% at ~$0.24, a 17x cost gap. Distillation mostly removed first-try reliability, not coverage—the pass@4 gap shrinks to 2.6 points. A Flash-then-escalate cascade reached 80.9% at about $1.70 per task, beating the full model alone at less than half the price. Coding agents should default to Flash and promote on test failure, but Flash broke an already-passing baseline in 6.9% of rollouts versus 4.4%, so diffs still need a regression gate.
Key Takeaways
- ✓On DeepSWE, GLM-5.3 is 69.0% pass@1 at ~$3.99; Flash is 63.4% at ~$0.24, about 17x cheaper.
- ✓A Flash-first cascade that escalates on test failure hits 80.9% at ~$1.70 per task.
- ✓Flash breaks passing baselines more often (6.9% vs 4.4%), so accept diffs only behind regression tests.