← Back to explorer

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-11T21:52:01+00:00
Verified
2026-08-11T21:52:01+00:00

Summary

Technical report on Qwen-Audio-3.0-TTS, a production-oriented text-to-speech system that the authors position as jointly advancing content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency and robustness. It pairs a 12.5 Hz low-frame-rate speech tokenizer (for reduced inference latency) with a five-stage progressive training paradigm that coordinates optimisation of the language model and the flow-matching model. Control is exposed at inference through free-style natural-language instructions plus fine-grained inline tags, and the system supports 16 languages and 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and generation from noisy, reverberant or unclear reference speech. Evaluated on SEED-TTS-Eval, CV3-Eval, instruction-following, long-form and acoustic-robustness suites, where the paper reports state-of-the-art or strongest aggregate results on many dimensions and first place on the independent Artificial Analysis Text-to-Speech Leaderboard.

Keywords

paper arxiv eess.as tts speech-synthesis text-to-speech multilingual qwen

Topics

Electrical Engineering and Systems Science / Audio and Speech Processing

Research notes

  • Posted in #random-papers as the HTML view (arxiv.org/html/2607.23938v1); the abs page is recorded here as canonical. Single version v1, submitted 27 Jul 2026, 19 pages, primary subject eess.AS. Benchmark claims (SOTA on many reported dimensions; first on the Artificial Analysis TTS leaderboard) are the authors’ own and were not independently verified. No dataset, code or model weights linked from the arXiv record.