Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, its first real-time audio perception model: streaming ASR, diarization for 20+ speakers, and endpointing, with native code-switching, 70+ training languages (25 validated at launch), and hour-plus sessions. Meta says it ranks first on Artificial Analysis streaming speech-to-text as of September 1. It already powers Meta AI for Mac (hold Fn to dictate anywhere) and Muse Code, and is on the Meta Model API at $3 per 1,000 audio-minutes (~$0.18/hour).
Key Takeaways
- βStreaming ASR, diarization, and endpointing ship as one model without a post-process chain
- βAlready live for system-wide Mac dictation and Muse Code; developers use the Meta Model API
- βAPI pricing is about $0.18/hour; adaptive delay trades latency for harder words