๐คฏ Future Transcription? Muse Voice Transcribe ๐
September 02, 2026 | Author ABR-INSIGHTS Tech Hub
Tech
๐ง Audio Summaries
๐ Shop on Amazon
ABR-INSIGHTS Tech Hub Picks
BROWSE COLLECTION โ*As an Amazon Associate, I earn from qualifying purchases.
Verified Recommendations๐ง Quick Intel
๐Summary
This week, Meta Superintelligence Labs introduced Muse Voice Transcribe, a new model from the Muse Spark family designed for advanced speech processing. The system combines transcription, speaker separation, and endpoint detection within a single, streaming architecture. It processes audio in 80ms chunks, utilizing soft tokens and predictive modeling with a delay reward. The model was trained on over 70 languages, achieving top performance on Artificial Analysis benchmarks for streaming speech-to-text and diarization, demonstrating a final WER of 3.1% at 0.16 seconds after speech. Meta reports a diarization error rate of 17.5% across several datasets, with pricing at $0.18 per hour, undercutting competitors. The technologyโs native support for long-context audio and multi-speaker environments represents a significant advancement in real-time speech processing capabilities.
๐กInsights
โผ
MUSE VOICE TRANSCRRIBE: A REVOLUTIONARY AUDIO PERCEPTION MODEL
Meta Superintelligence Labs has unveiled Muse Voice Transcribe, a groundbreaking autoregressive multimodal model designed to revolutionize real-time audio perception. This innovative system collapses traditionally separate processes โ transcription, speaker separation, and endpoint detection โ into a single, streamlined model, offering unprecedented speed and efficiency.
THE CORE ARCHITECTURE AND FUNCTIONALITY
Muse Voice Transcribe operates by processing audio in 80ms chunks at a 12.5 Hz sampling rate, transforming each chunk into a soft token. The model then makes a binary decision: either predict the next audio chunk (<|next_audio|> token) and continue listening, or emit a text token. This autoregressive approach, combined with a carefully managed โdelayโ โ the gap between listening and writing โ allows for a dynamic trade-off between accuracy and latency. The modelโs reinforcement learning training incorporates both word error rate and delay rewards, resulting in a policy that adapts delay based on the complexity of the spoken content.
STREAMING ASR, SPEAKER DIARIZATION, AND ENDPOINTING
The systemโs capabilities extend beyond basic transcription. It performs streaming ASR, speaker diarization for up to 20 speakers simultaneously, and endpoint detection, all within a single pass without requiring post-processing. Speaker diarization utilizes special tokens (<|start_of_turn|> token) to mark potential speaker switches and <|speaker_{A-Z}|> tags to identify individual speakers. The turn token triggers immediately upon a switch, while the speaker tag is delayed to the end of the chunk, allowing for flexible handling of overlapping speech. The model also supports native code-switching, a critical feature for bilingual speakers.
JOINT TRAINING AND MULTILINGUAL SUPPORT
Muse Voice Transcribe was trained on over 70 languages, with 25 extensively verified and recommended for launch. Crucially, the transcription, speaker diarization, and endpointing tasks are trained jointly, optimizing the overall performance. This approach leverages extra rewards layered on top of the ASR reward to fine-tune the modelโs behavior. The model supports long-context handling, natively supporting audio input exceeding one hour and 20+ speakers without requiring post-processing.
PERFORMANCE BENCHMARKS AND COMPETITIVE ANALYSIS
Independent analysis confirms Muse Voice Transcribeโs impressive performance. On Artificial Analysis AA-WER Streaming, it achieves a final-transcript WER of 3.1% at 0.16s after end of speech, surpassing competitors like Cartesia Ink-2 (3.4% at 0.43s) and ElevenLabs Scribe v2 Realtime (3.6% at 0.14s). In initial partial transcript evaluations, Muse Voice Transcribe records a WER of 3.6% at 0.13s. On diarization benchmarks (AMI-IHM, AMI-SDM, and VoxConverse), the model demonstrates an average diarization error rate of 17.5%, significantly lower than other systems ranging from 21.1% to 28.6%.
PRICING AND AVAILABILITY
Muse Voice Transcribe is available as a hosted API at $3.00 per 1,000 audio minutes ($0.18 per hour). This competitive pricing undercuts existing solutions like Cartesia Ink-2 ($4.00 per 1,000 minutes) and ElevenLabs Scribe v2 Realtime ($6.50 per 1,000 minutes). Currently, access is limited to Meta AI for Mac and Muse Code, with no publicly released model weights, restricting self-hosting options.
KEY TECHNICAL SPECIFICATIONS
The modelโs architecture leverages an autoregressive multimodal approach, processing audio chunks of 80ms at 12.5 Hz. The delay mechanism dynamically adjusts based on word error rate and latency considerations. Speaker diarization utilizes <|start_of_turn|> and <|speaker_{A-Z}|> tokens, while endpointing employs <|speech_onset|> and <|speech_endpoint|> markers. The system is trained on 70+ languages with native code-switching capabilities. It natively handles audio input exceeding one hour and 20+ speakers.
METAS REPORTED ACHIEVEMENTS
Meta reports that Muse Voice Transcribe places first on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks, as of September 1, 2026. This achievement highlights the modelโs advanced capabilities and positions it as a leading solution in the field of real-time audio perception.
Related Articles
Tech
AI Law Crisis ๐จ: Will ChatGPT Survive? ๐ค
The European Commission announced Monday that OpenAIโs ChatGPT, along with Reddit and Roblox, would be subject to strict...
Tech
Apple Shakes Off Cook ๐๐ฅ: Big Changes Ahead!
Tim Cook is departing his role as Appleโs CEO, triggering a series of executive transitions. Phil Schiller, head of the...
Tech
Robots Taking Data Centers ๐ค๐ฅ: Human Jobs at Risk?
Meta is currently testing robots within its data centers, utilizing machines from companies like Watney Robotics, Kinova...