HuggingPapers reports NVIDIA published NVFP4-quantized Qwen3.8-Flash-Next weights on Hugging Face (125B MoE with hybrid attention), claiming ~63% smaller size with minimal accuracy loss. The official repo nvidia/Qwen3.8-Flash-Next-NVFP4 targets higher-throughput local/edge inference on NVIDIA hardware with a lower VRAM floor.
Key Takeaways
- ✓Weights: nvidia/Qwen3.8-Flash-Next-NVFP4 is on Hugging Face.
- ✓Compression: NVFP4 cuts size ~63% for higher-throughput local inference.
- ✓Fit: 125B MoE + hybrid attention for coding/chat workloads on NVIDIA GPUs.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.