Together AI launched hosted serverless API serving for GLM-5.3 Flash, highlighting its 320B-A18B multimodal MoE architecture, 1M context, and high DeepSWE efficiency compared to proprietary frontier models.
Key Takeaways
- ✓Brings 320B total / 18B active multimodal MoE into hosted serverless APIs with instant elastic scaling;
- ✓Native visual reasoning combined with hybrid attention reduces input token consumption for agentic workflows;
- ✓Matches frontier-level engineering accuracy on DeepSWE while delivering more than double the completed tasks per dollar.