Alibaba Qwen team open-sources Qwen3.8-Flash-Next weights (125B total / 6B active MoE). Partnering with Unsloth, optimized GGUF quantizations run full 262K context on single 24GB GPUs at 85 tokens/sec.
Key Takeaways
- ✓Qwen team open-sources Qwen3.8-Flash-Next 125B/6B MoE weights;
- ✓Unsloth GGUF quantizations enable full 262K context on single 24GB GPUs;
- ✓Local inference throughput reaches 85 tokens/s for air-gapped private development.