Unsloth released 3-bit GGUF quantization for Z.ai's GLM-5.3-Flash (Ox Alpha), enabling local on-device inference for the 320B-A18B multimodal MoE on 128GB RAM workstations without datacenter GPUs.

Key Takeaways

  • Day-one local offline support for a 320B multimodal coding model on workstation hardware like Mac Studio;
  • Maintains high reasoning and coding fidelity on DeepSWE benchmarks under optimized 3-bit quantization;
  • Weights and deployment guides published on Hugging Face for instant use with llama.cpp and LM Studio.