Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, its first real-time audio perception model: streaming ASR, diarization for 20+ speakers, and endpointing, with native code-switching, 70+ training languages (25 validated at launch), and hour-plus sessions. Meta says it ranks first on Artificial Analysis streaming speech-to-text as of September 1. It already powers Meta AI for Mac (hold Fn to dictate anywhere) and Muse Code, and is on the Meta Model API at $3 per 1,000 audio-minutes (~$0.18/hour).

Key Takeaways

  • βœ“Streaming ASR, diarization, and endpointing ship as one model without a post-process chain
  • βœ“Already live for system-wide Mac dictation and Muse Code; developers use the Meta Model API
  • βœ“API pricing is about $0.18/hour; adaptive delay trades latency for harder words
ADSponsored