Tech

Meta Superintelligence Lab launches Muse Voice Transcribe for real-time multilingual audio

The new model distinguishes between more than 20 speakers and 70 languages simultaneously, marking a significant step in real-time audio perception for developers and Mac users.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Engadget · View original source
Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time
Technology

Meta Superintelligence Lab has released Muse Voice Transcribe, its first real-time audio perception model designed to handle complex dictation and transcription tasks. The system is capable of managing more than 20 speakers while seamlessly processing multiple languages at once, a capability that addresses the challenges of messy, real-world audio environments.

The model natively handles speaker diarisation and endpointing within a single architecture, eliminating the need for separate processing layers. According to Meta, the system uses adaptive delay to predict tokens, allowing it to wait longer on difficult words and commit faster on easier ones to increase overall accuracy. This approach enables the model to manage hour-long sessions effectively.

Language support is a key feature of the release, with the model trained across more than 70 languages. While 25 languages have been validated at launch, the system is designed to handle mid-sentence code-switching, where speakers alternate between languages within a single sentence. This functionality is particularly relevant for global audiences and multilingual workplaces.

Muse Voice Transcribe is currently available in the Meta AI Mac app, Muse Code, and via Meta’s Model API. Because the Meta AI Mac app powers voice-enabled features on other applications, the new model will now drive dictation features on third-party services. For developers, the API is priced at $3 for 1,000 audio minutes, with a demo version also available on Meta’s research blog.

The launch comes less than a week after Google released Gemini 3.5 Transcribe, an audio model with similar capabilities that is being integrated into Android and Chrome. While Google is embedding its technology into its flagship operating systems, it remains unclear whether Meta plans to integrate Muse Voice Transcribe into its own core services beyond the Mac app and developer tools.

Meta CEO Mark Zuckerberg, who recently returned to the X platform after a three-year hiatus, shared an example of the model’s capabilities, highlighting its ability to distinguish between multiple speakers and switch languages automatically. The release is part of a broader push by Meta Superintelligence Lab, which has recently introduced a dedicated coding agent, an open-weight model, and the Meta AI Mac app.

Continue reading

More from Tech

Read next: Tesla sets 10 October reveal for long-delayed second-generation Roadster
Read next: Trump and Johnson reject calls to slow frontier AI development
Read next: AI doomer warnings put Anthropic’s business interests under scrutiny