MiaAI Lab (@MiaAI_lab) added a TP=3 path to GLM-5.3-Flash EXL3 for three DGX Sparks: +28% decode and +9% prefill vs TP=2, with 3.2M KV cache and 1M context, plus optional NFS weight sharing from the head node (NFS_SHARE=1). Repo: github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks.
Key Takeaways
- βNew TP=3 path runs GLM-5.3-Flash EXL3 across three DGX Sparks.
- βVs TP=2: +28% decode, +9% prefill; 3.2M KV cache and 1M context.
- βOptional NFS_SHARE=1 loads weights from the head node to free other Sparks.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.