Alibaba Qwen confirmed that the 125B multimodal MoE Qwen3.8-Flash-Next can now run locally on about 75GB of RAM via Unsloth GGUFs, with N-gram embeddings offloaded to CPU or unified memory for near-VRAM speed. It is the day-0 local path for the Qwen4 architecture preview weights.

Key Takeaways

  • โœ“125B MoE with 6B active parameters plus 51B n-gram embeddings now runs locally on ~75GB RAM without fitting the full model in GPU VRAM
  • โœ“Unsloth shipped day-0 GGUFs and a Qwen3.8-Flash-Next guide; CPU RAM / unified-memory setups approach VRAM throughput
  • โœ“Official weights remain on Hugging Face and ModelScope, with day-0 vLLM and SGLang serving on NVIDIA and AMD