Qwen-Music Technical Report
- Type
- paper
- Venue
- arXiv / Qwen Team
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:32:00Z
- Verified
- 2026-08-14T16:32:00Z
Summary
Three-stage system: 25 Hz single-codebook Music Semantic Tokenizer (0.6B Conformer), Qwen-Music-LLM (3B dense Qwen3.5-Omni init with Melody-CoT planning), and Qwen-Music-Render (1.3B DiT + Spec-VAE + Band-Mode Refiner to 48 kHz stereo). LLM trained on >5M hours of multilingual music with a quality-graded curriculum then SFT/DPO/GSPO. On 600 ZH/EN prompts, best on 13/16 SongBench/SongEval/AudioBox-Aesthetic metrics; professional A/B prefers it over MiniMax Music 2.6 (66.7%) and Suno V5 (55.4%), comparable to Suno V5.5 (50.3%). Cover mode reports lower Melody MAE than Suno V5.5/V5 and MiniMax Cover on an AI-generated reference set.
Keywords
qwen-music · text-to-music · cover-song · melody-cot · tokenizer · dit · songbench · suno
Topics
music generation, audio, song synthesis
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.SD). Discord #data-source-dump posted AlphaXiv 2607.11699v1. SOTA and A/B claims are the authors'. External AA Music-with-Vocals listing as JazzCat was not re-checked. No official weights/code linked from the abs page; did not copy an Apache-2.0 claim from secondary writeups.