Together AI launched hosted Day-1 inference for GLM-5.3 Flash (320B-A18B MoE, 1M context), nearly matching Luna on DeepSWE benchmarks while completing more than twice the work per dollar.
Key Takeaways
- ✓Day-1 managed serving of 320B-A18B MoE lowers deployment hurdles for custom hybrid-attention architectures;
- ✓DeepSWE evaluations demonstrate over 2x work completed per dollar compared to closed frontier models;
- ✓Available immediately via Together AI Model Market and unified developer API for dedicated endpoints.