Baseten launched GLM-5.3-Flash on its Model APIs platform on day zero, achieving 122+ TPS with default Zero Data Retention (ZDR) and US-only infrastructure at a 90% cost reduction over GLM-5.2.
Key Takeaways
- ✓Achieves top-tier 122+ TPS decode throughput across Artificial Analysis and OpenRouter independent benchmarks;
- ✓Enterprise compliance guaranteed via US-only isolated compute clusters and strict Zero Data Retention (ZDR);
- ✓1M context and native multimodal vision make it a prime cost-effective engine for enterprise coding agents.