← Back to explorer

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling Up Real-world Acoustic Simulation

Type
paper
Venue
arXiv / NTU / NUS / Shanghai AI Lab
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T20:26:00Z
Verified
2026-08-14T20:26:00Z

Summary

NTU/NUS/Shanghai AI Lab (cite Xie et al.; arXiv 2605.19833). Voices-in-the-Wild-2M: 7 atomic effects (noise, far-field, obstructed, echo&reverb, recording, electronic distortion, transmission dropout) composed into 54 physically plausible scenarios; linear severity; drop WER>70% samples. A2S-SFT on Qwen3-ASR-1.7B (encoder/aligner curriculum then LLM then joint) plus Dual-Granularity WER-Gated Policy Optimization (token-level refinement vs sentence-level reconstruction, gate at WER 0.3). Mega-ASR 6.70 avg WER on CHiME-4/VOiCES/NOIZEUS vs Qwen3-ASR 7.93; VOiCES rm4 far babble 45.69 vs 54.01; NOIZEUS 0dB 19.80 vs 23.97; mixed Voices-in-the-Wild-Bench 2.73/4.57 vs Qwen3-ASR 3.30/5.39. Environment-aware LoRA router preserves clean ASR. Introduces Voices-in-the-Wild-Bench (5k clips).

Keywords

mega-asr · robust-asr · voices-in-the-wild · acoustic-simulation · qwen3-asr · ntu

Topics

robust ASR, acoustic simulation, audio-language models

Research notes

  • Primary: arxiv abs 2605.19833 (cs.SD). Project https://xzf-thu.github.io/Mega-ASR/. Code https://github.com/xzf-thu/Mega-ASR (no SPDX). Dataset https://huggingface.co/datasets/zhifeixie/Voices-in-the-Wild-2M. Discord posted abs + project page. Substantial speech corpus — papers_local intake only; do not write datasets_local here. License left blank (no CC on abs, no GitHub SPDX).