← Back to explorer

Qwen-Music Technical Report

Type
paper
Venue
arXiv / Qwen Team
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:32:00Z
Verified
2026-08-14T16:32:00Z

Summary

Three-stage system: 25 Hz single-codebook Music Semantic Tokenizer (0.6B Conformer), Qwen-Music-LLM (3B dense Qwen3.5-Omni init with Melody-CoT planning), and Qwen-Music-Render (1.3B DiT + Spec-VAE + Band-Mode Refiner to 48 kHz stereo). LLM trained on >5M hours of multilingual music with a quality-graded curriculum then SFT/DPO/GSPO. On 600 ZH/EN prompts, best on 13/16 SongBench/SongEval/AudioBox-Aesthetic metrics; professional A/B prefers it over MiniMax Music 2.6 (66.7%) and Suno V5 (55.4%), comparable to Suno V5.5 (50.3%). Cover mode reports lower Melody MAE than Suno V5.5/V5 and MiniMax Cover on an AI-generated reference set.

Keywords

qwen-music · text-to-music · cover-song · melody-cot · tokenizer · dit · songbench · suno

Topics

music generation, audio, song synthesis

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.SD). Discord #data-source-dump posted AlphaXiv 2607.11699v1. SOTA and A/B claims are the authors'. External AA Music-with-Vocals listing as JazzCat was not re-checked. No official weights/code linked from the abs page; did not copy an Apache-2.0 claim from secondary writeups.