Muse Voice Transcribe
Meta's first real-time audio perception model. It does streaming ASR, speaker diarization and endpointing in 70+ languages.
Streaming speech-to-text at 3.1% word error rate and diarization at 17.5% error for 20+ speakers (Meta numbers); ships in Meta Model API, Meta AI for Mac and Muse Code.
- Date
- Tuesday, 1 September 2026
- Lab
- Meta
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Streaming word error rate | 3.1% Meta says first on Artificial Analysis; diarization error 17.5% | company |
Rank claim is Meta-reported.
Sources
This record was checked against its sources on 6 October 2026. How we check