🤯Real-Time AI Translation: Qwen3.8 Unlocks 🚀

September 20, 2026 |

AI

🎧 Audio Summaries
English flag
French flag
German flag
Japanese flag
Korean flag
Mandarin flag
Spanish flag

🧠Quick Intel


  • Qwen released Qwen3.8-LiveTranslate, a next-generation real-time simultaneous interpretation model utilizing a new Interleave architecture.
  • Average Lagging Audio Loss (LAAL) decreased from 2.8 seconds to 2.3 seconds, representing a significant improvement in translation speed.
  • The model supports a 53,248-token context window (49,152 input, 4,096 output) enabling longer-context disambiguation.
  • Real-time speaker diarization and synchronized bilingual display were added to the model’s capabilities.
  • Deployment is available via a hosted API on Alibaba Cloud Model Studio and QwenCloud as qwen3.8-livetranslate-flash-realtime over WebSocket.
  • Default rate limits are 10 requests and 100,000 tokens per minute.
  • Singapore pricing is $5.653 per 1M tokens for audio, with additional costs for text and image tokens.
  • 📝Summary


    Qwen has introduced Qwen3.8-LiveTranslate, a new model designed for real-time simultaneous interpretation. The system processes live speech, potentially with video, generating translated text and speech concurrently. A key advancement is the Interleave architecture, resulting in improvements in faithfulness, fluency, and conciseness, with a reduction in average lagging from 2.8 to 2.3 seconds. The release incorporates speaker diarization, bilingual displays, and long-context disambiguation, deployable via a hosted API on platforms like Alibaba Cloud. The model utilizes a 53,248-token context window and operates with default rate limits. Pricing in Singapore is approximately $5.653 per million tokens, with audio consumption rates of 7 and 12.5 tokens per second respectively. This technology represents a significant step in facilitating immediate translation across various applications.

    💡Insights



    Qwen3.8-LiveTranslate: A Revolution in Real-Time Interpretation
    Qwen has unveiled Qwen3.8-LiveTranslate, a groundbreaking real-time simultaneous interpretation model designed for unparalleled accuracy and responsiveness. This innovative system leverages a novel Interleave architecture to deliver translated text and speech synchronously with the speaker, effectively eliminating the delay inherent in traditional translation methods. Initial testing has demonstrated significant improvements across key metrics, including faithfulness, fluency, and conciseness, with a remarkable reduction in Length-Adaptive Average Lagging (LAAL) from 2.8 seconds to a highly competitive 2.3 seconds. This advancement represents a substantial leap forward in real-time translation capabilities, making it suitable for a wide range of applications demanding immediate and accurate communication. The core functionality centers around listening to live speech, optionally incorporating video frames, and generating translated output – a process meticulously optimized through the Interleave architecture.

    Key Features and Technical Specifications
    The Qwen3.8-LiveTranslate model incorporates several sophisticated features beyond its core translation capabilities. Notably, it includes real-time speaker diarization, allowing for the identification and labeling of different speakers during the interpretation process. A synchronized bilingual display provides immediate translation alongside the original speech, enhancing comprehension. Furthermore, the model incorporates long-context disambiguation, mitigating potential misunderstandings arising from ambiguous phrases or complex terminology. Deployable via a hosted API, the model is currently available on Alibaba Cloud Model Studio and QwenCloud, accessible through the qwen3.8-livetranslate-flash-realtimeover WebSocket interface. The system’s performance is directly measured by LAAL, a metric that quantifies the average delay between the source speech and the translated output. This targeted reduction of 18% in average lag highlights the model’s efficiency and responsiveness. The QwenCloud team positions this model as the real-time counterpart to Qwen3.8-LiveTranslate-Flash, built upon the robust Qwen-Omni stack, which utilizes large-scale multimodal data, cross-language and cross-modal alignment, and visual enhancement techniques. This architecture supports offline audio and video translation, expanding the model’s versatility. The model supports 60 languages, with 29 languages offering both audio and text output, and 31 languages providing only text output, including Chinese, English, Arabic, German, French, Spanish, Japanese, Korean, Hindi, and many others.

    Pricing, Usage, and Support
    The pricing structure for Qwen3.8-LiveTranslate is tiered, reflecting the varying demands of different usage scenarios. Singapore pricing stands at $5.653 per 1 million tokens, while Beijing offers more competitive rates of $0.466, $14.133, and $22.613 USD per 1 million tokens. The model's efficiency is quantified by token consumption: audio input consumes 7 tokens per second, and audio output consumes 12.5 tokens per second. A one-hour session, including both audio and text tokens, costs approximately $1.54 in Singapore. The context window is substantial, boasting 53,248 tokens with 49,152 allocated for input and 4,096 for output. Default rate limits are set at 10 requests and 100,000 tokens per minute. It’s important to note that certain functionalities, such as function calling, structured outputs, batch inference, and fine-tuning, are currently unsupported by the model. Detailed technical specifications can be found in the Technical Details section. Finally, Qwen encourages user engagement through various channels, including Twitter, a 150k+ member ML SubReddit, and a newsletter subscription. For partnership opportunities related to promoting GitHub repositories, Hugging Face pages, product releases, or webinars, users can connect directly through a dedicated platform boasting over 2 million monthly views.