Alibaba Qwen open-sourced Qwen3.8-Flash, a multimodal MoE with 125B parameters plus 51B n-gram embeddings and only 6B active per token. It scores 58.7 on DeepSWE 1.1 and 62.5 on SWE-bench Pro, was trained at 1/9 the cost of Qwen3.7-Plus, and ships Qwen3.8-Flash-Next as an early look at the Qwen4 architecture.
Key Takeaways
- โScores 58.7 DeepSWE 1.1, 62.5 SWE-bench Pro, and 73.9 CoWorkBench, beating Qwen3.7-Plus especially on coding and office work
- โGDN+QSA hybrid attention, Gated Residual, n-gram embeddings, and Muon; QSA is up to 7.6x faster in 1M-token prefill
- โWeights on Hugging Face and ModelScope with day-0 vLLM/SGLang support; QwenCloud API at $0.16/$0.47 per 1M tokens