Meta overhauled the inference engine for Muse Spark 1.3, combining sub-quadratic linear attention with lossless parallel diffusion decoding. The upgrade delivers a 3.8x throughput increase and 50% memory reduction during long-horizon agent coding workflows.

Key Takeaways

  • Pairs linear attention with diffusion decoding to achieve 3.8x generation throughput
  • Cuts VRAM footprints by half, enabling massive context generation on consumer GPUs
  • Cements Meta open-weights dominance across edge and high-concurrency developer systems
ADSponsored