๐Ÿคฏ Future Transcription? Muse Voice Transcribe ๐Ÿš€

September 02, 2026 |

Tech

๐ŸŽง Audio Summaries
English flag
French flag
German flag
Japanese flag
Korean flag
Mandarin flag
Spanish flag
๐Ÿ›’ Shop on Amazon

๐Ÿง Quick Intel


  • Meta Superintelligence Labs launched Muse Voice Transcribe, an autoregressive multimodal model from the Muse Spark family, offering streaming ASR, speaker diarization, and endpoint detection.
  • The model processes 80ms audio chunks at 12.5 Hz, utilizing a soft token approach and a single decoder loop for audio context and delay control.
  • Reinforcement learning combines a word error rate reward with a delay reward, optimizing for both accuracy and latency.
  • Muse Voice Transcribe is available via a hosted API at $3.00 per 1,000 audio minutes ($0.18 per hour), powering dictation in Meta AI for Mac and Muse Code.
  • The model was trained on 70+ languages, with 25 extensively verified and recommended at launch, supporting native code-switching.
  • On Artificial Analysis AA-WER Streaming, Muse Voice Transcribe achieves 3.1% final-transcript WER at 0.16s after end of speech, surpassing Cartesia Ink-2 (3.4% at 0.43s) and ElevenLabs Scribe v2 Realtime (3.6% at 0.14s).
  • Meta reports a 17.5% average diarization error rate across AMI-IHM, AMI-SDM, and VoxConverse, significantly better than other systems ranging from 21.1% to 28.6%.
  • As of September 1, 2026, Muse Voice Transcribe achieved first place on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks.
  • ๐Ÿ“Summary


    This week, Meta Superintelligence Labs introduced Muse Voice Transcribe, a new model from the Muse Spark family designed for advanced speech processing. The system combines transcription, speaker separation, and endpoint detection within a single, streaming architecture. It processes audio in 80ms chunks, utilizing soft tokens and predictive modeling with a delay reward. The model was trained on over 70 languages, achieving top performance on Artificial Analysis benchmarks for streaming speech-to-text and diarization, demonstrating a final WER of 3.1% at 0.16 seconds after speech. Meta reports a diarization error rate of 17.5% across several datasets, with pricing at $0.18 per hour, undercutting competitors. The technologyโ€™s native support for long-context audio and multi-speaker environments represents a significant advancement in real-time speech processing capabilities.

    ๐Ÿ’กInsights

    โ–ผ


    MUSE VOICE TRANSCRRIBE: A REVOLUTIONARY AUDIO PERCEPTION MODEL
    Meta Superintelligence Labs has unveiled Muse Voice Transcribe, a groundbreaking autoregressive multimodal model designed to revolutionize real-time audio perception. This innovative system collapses traditionally separate processes โ€“ transcription, speaker separation, and endpoint detection โ€“ into a single, streamlined model, offering unprecedented speed and efficiency.

    THE CORE ARCHITECTURE AND FUNCTIONALITY
    Muse Voice Transcribe operates by processing audio in 80ms chunks at a 12.5 Hz sampling rate, transforming each chunk into a soft token. The model then makes a binary decision: either predict the next audio chunk (<|next_audio|> token) and continue listening, or emit a text token. This autoregressive approach, combined with a carefully managed โ€œdelayโ€ โ€“ the gap between listening and writing โ€“ allows for a dynamic trade-off between accuracy and latency. The modelโ€™s reinforcement learning training incorporates both word error rate and delay rewards, resulting in a policy that adapts delay based on the complexity of the spoken content.

    STREAMING ASR, SPEAKER DIARIZATION, AND ENDPOINTING
    The systemโ€™s capabilities extend beyond basic transcription. It performs streaming ASR, speaker diarization for up to 20 speakers simultaneously, and endpoint detection, all within a single pass without requiring post-processing. Speaker diarization utilizes special tokens (<|start_of_turn|> token) to mark potential speaker switches and <|speaker_{A-Z}|> tags to identify individual speakers. The turn token triggers immediately upon a switch, while the speaker tag is delayed to the end of the chunk, allowing for flexible handling of overlapping speech. The model also supports native code-switching, a critical feature for bilingual speakers.

    JOINT TRAINING AND MULTILINGUAL SUPPORT
    Muse Voice Transcribe was trained on over 70 languages, with 25 extensively verified and recommended for launch. Crucially, the transcription, speaker diarization, and endpointing tasks are trained jointly, optimizing the overall performance. This approach leverages extra rewards layered on top of the ASR reward to fine-tune the modelโ€™s behavior. The model supports long-context handling, natively supporting audio input exceeding one hour and 20+ speakers without requiring post-processing.

    PERFORMANCE BENCHMARKS AND COMPETITIVE ANALYSIS
    Independent analysis confirms Muse Voice Transcribeโ€™s impressive performance. On Artificial Analysis AA-WER Streaming, it achieves a final-transcript WER of 3.1% at 0.16s after end of speech, surpassing competitors like Cartesia Ink-2 (3.4% at 0.43s) and ElevenLabs Scribe v2 Realtime (3.6% at 0.14s). In initial partial transcript evaluations, Muse Voice Transcribe records a WER of 3.6% at 0.13s. On diarization benchmarks (AMI-IHM, AMI-SDM, and VoxConverse), the model demonstrates an average diarization error rate of 17.5%, significantly lower than other systems ranging from 21.1% to 28.6%.

    PRICING AND AVAILABILITY
    Muse Voice Transcribe is available as a hosted API at $3.00 per 1,000 audio minutes ($0.18 per hour). This competitive pricing undercuts existing solutions like Cartesia Ink-2 ($4.00 per 1,000 minutes) and ElevenLabs Scribe v2 Realtime ($6.50 per 1,000 minutes). Currently, access is limited to Meta AI for Mac and Muse Code, with no publicly released model weights, restricting self-hosting options.

    KEY TECHNICAL SPECIFICATIONS
    The modelโ€™s architecture leverages an autoregressive multimodal approach, processing audio chunks of 80ms at 12.5 Hz. The delay mechanism dynamically adjusts based on word error rate and latency considerations. Speaker diarization utilizes <|start_of_turn|> and <|speaker_{A-Z}|> tokens, while endpointing employs <|speech_onset|> and <|speech_endpoint|> markers. The system is trained on 70+ languages with native code-switching capabilities. It natively handles audio input exceeding one hour and 20+ speakers.

    METAS REPORTED ACHIEVEMENTS
    Meta reports that Muse Voice Transcribe places first on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks, as of September 1, 2026. This achievement highlights the modelโ€™s advanced capabilities and positions it as a leading solution in the field of real-time audio perception.