← Back to explorer

Audio-FLAN Dataset

Type
dataset
Venue
HKUST Audio (Hong Kong University of Science and Technology)
Year
2026
Source
huggingface
Access
restricted
Language
English, Chinese
Added
2026-07-17T20:18:03.731422+00:00
Verified
2026-07-17T20:18:03.731422+00:00

Summary

Audio-FLAN is a large-scale instruction-tuning dataset for unified audio understanding and generation, integrating 50+ source datasets across speech, music, and general audio domains. It covers tasks like speech recognition, text-to-speech, acoustic scene classification, audio event recognition, and music generation, organized with structured metadata (instruction, input, output, task type) for training audio-language models.

Keywords

audio speech music sound instruction-tuning text-to-speech asr audio-generation

Topics

Audio / Speech / Music

Research notes

  • Requires sharing contact information for access. Audio included only for datasets with permissive licenses (CC-BY, Apache-2.0, MIT); metadata-only for restrictive licenses. Paper: arxiv 2502.16584.