Z.ai launched GLM-5.3-Flash, a natively multimodal 320B-A18B MIT-licensed model with a 1M-token context, previously previewed as the stealth Ox Alpha. On Z.ai's Code Bench it matches Claude Opus 4.8 at $0.15/$0.50 per million input/output tokens, and it runs entirely on Chinese AI chips.

Key Takeaways

  • MIT-licensed 320B-A18B sparse model with native multimodality and a 1M-token context
  • Matches Claude Opus 4.8 on Z.ai's real-world Code Bench at roughly an order-of-magnitude lower price
  • Previously the stealth Ox Alpha on OpenRouter; weights, API, ZCode, and Coding Plan are live
ADSponsored