← Back to explorer

Qwen2.5-Omni Technical Report

Type
other
Venue
arXiv / Qwen Team

Summary

Block-wise audio/vision encoders feed a shared LLM; TMRoPE interleaves time-aligned audio/video. Thinker generates text; Talker is a dual-track AR speech head on Thinker hidden states plus sliding-window DiT. Qwen2.5-Omni-7B sits between Qwen2-7B and Qwen2.5-7B on text (MMLU-redux 71.0 vs Qwen2.5-7B 75.4 / Qwen2-7B 67.3). OmniBench avg 56.13 vs Gemini-1.5-Pro 42.91; MMMU val 59.2; Video-MME w/o sub 64.3; speech-instruction MMLU 65.6 vs text Qwen2-7B 69.3. Streaming TTS WER 1.42/2.33/6.54 on SEED test-zh/en/hard. Code https://github.com/QwenLM/Qwen2.5-Omni; blog https://qwenlm.github.io/blog/qwen2.5-omni/.

Keywords

qwen2.5-omni · tmrope · thinker-talker · streaming · multimodal · asr · tts · omnibench

Topics

multimodal LLMs, streaming speech, TMRoPE

Research notes

  • Primary: arxiv abs (cs.CL; also cs.CV, cs.SD, eess.AS). License CC BY 4.0 on HTML at check. Authors listed as Qwen Team with 14 named authors on the abs plus a long alphabetical contributor list. Code https://github.com/QwenLM/Qwen2.5-Omni (4,065 stars at check). Project https://qwenlm.github.io/blog/qwen2.5-omni/. HF paper page 174 upvotes; githubRepo linked; official linked models Qwen/Qwen2.5-Omni-7B and Qwen/Qwen2.5-Omni-3B (plus unofficial GGUFs) not copied into hf_* fields. Unofficial linked dataset Archit00/bbench-dep-marble not substantial, so no datasets_local row. Discord posted PDF with a note on TMRoPE.