Alibaba released Qwen3.8-Flash, a multimodal MoE previewing the Qwen4 architecture. Featuring 125B total params with 6B active per token, it scores 62.5 on SWE-bench Pro and 58.7 on DeepSWE 1.1. Trained at 1/9 the cost of Qwen3.7-Plus, it lists on QwenCloud at ~$0.16 in / $0.47 out per 1M tokens.
Key Takeaways
- โGDN+QSA hybrid attention with Muon optimizer drives SWE-bench Pro score to 62.5.
- โ125B total parameters with only 6B active per token deliver high inference speed and low cost.
- โNative 262k context window extends up to 1M tokens with YaRN.