Unsloth released 3-bit GGUF quantization for Z.ai's GLM-5.3-Flash (Ox Alpha), enabling local on-device inference for the 320B-A18B multimodal MoE on 128GB RAM workstations without datacenter GPUs.
Key Takeaways
- ✓Day-one local offline support for a 320B multimodal coding model on workstation hardware like Mac Studio;
- ✓Maintains high reasoning and coding fidelity on DeepSWE benchmarks under optimized 3-bit quantization;
- ✓Weights and deployment guides published on Hugging Face for instant use with llama.cpp and LM Studio.