Together AI launched hosted serverless API serving for GLM-5.3 Flash, highlighting its 320B-A18B multimodal MoE architecture, 1M context, and high DeepSWE efficiency compared to proprietary frontier models.

Key Takeaways

  • Brings 320B total / 18B active multimodal MoE into hosted serverless APIs with instant elastic scaling;
  • Native visual reasoning combined with hybrid attention reduces input token consumption for agentic workflows;
  • Matches frontier-level engineering accuracy on DeepSWE while delivering more than double the completed tasks per dollar.