MiaAI Lab (@MiaAI_lab) added a TP=3 path to GLM-5.3-Flash EXL3 for three DGX Sparks: +28% decode and +9% prefill vs TP=2, with 3.2M KV cache and 1M context, plus optional NFS weight sharing from the head node (NFS_SHARE=1). Repo: github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks.

Key Takeaways

  • βœ“New TP=3 path runs GLM-5.3-Flash EXL3 across three DGX Sparks.
  • βœ“Vs TP=2: +28% decode, +9% prefill; 3.2M KV cache and 1M context.
  • βœ“Optional NFS_SHARE=1 loads weights from the head node to free other Sparks.
ADSponsored