Alibaba Qwen team open-sources Qwen3.8-Flash-Next weights (125B total / 6B active MoE). Partnering with Unsloth, optimized GGUF quantizations run full 262K context on single 24GB GPUs at 85 tokens/sec.

Key Takeaways

  • Qwen team open-sources Qwen3.8-Flash-Next 125B/6B MoE weights;
  • Unsloth GGUF quantizations enable full 262K context on single 24GB GPUs;
  • Local inference throughput reaches 85 tokens/s for air-gapped private development.