Alibaba Qwen announced that the production Qwen3.8-Flash API is live on QwenCloud at $0.16 input and $0.47 output per 1M tokens. The multimodal MoE activates only 6B parameters per token and is positioned as cheaper and stronger than Qwen3.7-Plus on coding and office workloads.
Key Takeaways
- โProduction qwen3.8-flash is live on QwenCloud at $0.16 / $0.47 per 1M input/output tokens
- โ125B parameters plus 51B n-gram embeddings with 6B active per token; official scores are 58.7 DeepSWE 1.1 and 62.5 SWE-bench Pro
- โ262K native context expandable to 1M with YaRN; open-weight Qwen3.8-Flash-Next is already deployable locally and via vLLM/SGLang