Qwen3.8-LiveTranslate Release — Advancements in Real-Time Multimodal Speech Translation

Release of Qwen3.8-LiveTranslate

Alibaba has introduced Qwen3.8-LiveTranslate, a specialized model designed for real-time speech-to-speech translation that significantly improves response speed while maintaining vocal characteristics.

Key Technical Capabilities

  • Reduced Latency (Low Delay): The delay/latency certain tasks way down from 2.8 to 2.3 seconds.
  • Advanced Speaker Management: The model can distinguish between different speakers in multi-person conversations, assign labels to who is speaking, and attempt to preserve each person's individual voice timbre during outputted playback.
  • Multimodal Contextual Awareness: To improve accuracy regarding names, terms, and ambiguous phrases, the model utilizes visual context such as images, gestures, and text appearing on screen or previously spoken replies.
  • Broad Language Support: Supports any kind worth checking out up to 60 languages; however, it provides both text AND translated speech for 29 of those supported languages.

Access to this capability via API enough so dass developers check qwencloud if they want access today.
Demo available at omni.qwen.ai
Information source provided by dailyprompts.ru and MAX.

The new update marks an advancement toward more naturalized human conversation through low latency and speaker preservation.

! DYOR (Do Your Own Research)