Meta is launching Muse Voice Transcribe, its first real-time audio perception model, bringing multilingual streaming transcription to Mac, Muse Code, and developers via the Meta Model API. On the Meta AI Mac app, users can hold the Fn key to dictate into any application without switching contexts or relying on Apple’s native dictation.
Adaptive Delay Changes How Real-Time Dictation Works
Muse Voice Transcribe does not lock in a single speed-versus-accuracy tradeoff. Instead, it uses adaptive delay which means that the model listens briefly for straightforward words and commits them instantly, then waits longer when it encounters difficult speech or ambiguous phonemes. This approach mirrors how humans process conversation, faster for clear passages, slower for mumbled or noisy segments, and avoids the latency felt with traditional systems that use one fixed delay for every word.
The model combines streaming automatic speech recognition with speaker diarization and endpointing in a single pass. It can separate 20-plus voices across recordings, identify code-switching within or between sentences across more than 70 trained languages (25 validated at launch), and handle audio longer than an hour without post-processing. Language, keyword, and context biasing can further boost accuracy.
Pricing and Developer Access
Developers can access Muse Voice Transcribe through the Meta Model API at $3 per 1,000 audio-minutes, or $0.18 per hour. This undercuts several commercial speech-to-text competitors and signals Meta’s aggressive positioning in the developer API market.
End users using Meta AI for Mac or Muse Code face no additional charge; dictation is built in and active immediately. Hold Fn and speak into Mail, web browsers, code editors, or any text field to begin dictating today.
Performance Rankings and Validation
Meta claims Muse Voice Transcribe ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026. Independent benchmarking provides third-party validation against competitors, though results remain bound by the benchmark’s scope and methodology.
Muse Voice Fits Meta’s Rapid Model Expansion
This launch extends Meta Superintelligence Labs’ fast-moving roadmap under Alexandr Wang. Since April 2026, the lab has shipped Muse Spark (a reasoning model), Muse Code (a terminal-based AI coding agent for macOS and Linux), and now Muse Voice Transcribe. This shows that Meta has moved from its open-source Llama approach toward proprietary frontier models powering both consumer products and developer APIs across Meta’s platform ecosystem.