LightSeek Foundation released a Kimi K3 draft-model collection (EAGLE-3, DFlash2, DSpark) trained with TorchSpec on 40 NVIDIA GB200 GPUs, with open data recipes. Their blog reports acceptance length across 10 benchmarks and SPEED-Bench end-to-end throughput; DFlash2 averages ~4.02 accepted tokens and leads EAGLE-3 at lower concurrency.

Key Takeaways

  • Three draft architectures for Kimi K3 speculative decoding: EAGLE-3, DFlash2, DSpark.
  • Trained on 40×GB200 with TorchSpec; data recipes plus Docker/CI shared.
  • On HumanEval, DFlash2 acceptance length ~5.08 vs EAGLE-3 ~3.29.
ADSponsored