Meta has introduced Muse Voice Transcribe, its first real-time audio perception model from Meta Superintelligence Labs. The company announced the model on September 1, 2026. It brings streaming speech recognition, speaker identification and endpoint detection into one system.
More importantly, Muse Voice Transcribe already powers dictation in Meta AI for Mac and Muse Code. Developers can also access it through the Meta Model API. On Mac, users can hold the Fn key to dictate into applications across the system.
Real-Time Dictation Comes to Mac
Muse Voice Transcribe combines three core audio capabilities. First, it converts speech into text as users speak. Second, it identifies different speakers during longer recordings. Finally, it detects when a speaker starts or finishes talking.
The model also supports multilingual conversations and code-switching. Consequently, users can move between languages within the same sentence without switching transcription systems. Meta says the model is trained on more than 70 languages, with 25 languages extensively validated for the initial release.
For India, the launch includes native support for Hindi, Tamil, Telugu, Kannada and Malayalam. This makes Muse Voice Transcribe particularly relevant for multilingual voice interactions and mixed-language conversations.
The system can also process audio lasting more than one hour. Moreover, it can distinguish more than 20 speakers without requiring a separate post-processing stage.
Adaptive Delay Targets Speed and Accuracy
Meta has designed Muse Voice Transcribe to adjust its response timing according to speech difficulty. Instead of using one fixed delay, the model decides how long to listen before producing each word.
As a result, it can respond quickly to straightforward speech while waiting longer when additional context could improve accuracy. Meta calls this approach “adaptive delay.” The company uses reinforcement learning to balance transcription accuracy against latency.
Technically, Muse Voice Transcribe belongs to the Muse Spark family of autoregressive multimodal models. It processes incoming audio in 80-millisecond chunks. At each stage, the model decides whether to continue listening or produce a text token.
Meta also says the model ranks first on Artificial Analysis’ streaming speech-to-text leaderboard and on public diarization benchmarks.
Meta Opens Model to Developers
Beyond Mac dictation, Meta has made Muse Voice Transcribe available through its Model API. The service costs $3 per 1,000 audio minutes, according to current reports. That pricing gives developers another option for building real-time speech applications.
Meanwhile, Muse Code uses the same technology for voice-based interaction. Therefore, developers can use speech input while working with Meta’s coding assistant.
The broader strategy extends beyond transcription. Meta describes real-time audio perception as a foundation for more capable personal AI systems. The company says better speech recognition can help AI understand conversations, distinguish speakers and respond with greater context.
However, benchmark leadership does not by itself establish performance across every real-world environment. Independent testing will remain important, particularly for accents, noisy settings and less widely represented languages.
For now, Muse Voice Transcribe marks Meta’s move into real-time audio AI at the model level. With Mac dictation, Muse Code and API access available, the technology is already moving beyond a research demonstration into practical applications.








