Audio-FLAN Dataset
- Type
- dataset
- Venue
- HKUST Audio (Hong Kong University of Science and Technology)
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English, Chinese
- Added
- 2026-07-17T20:18:03.731422+00:00
- Verified
- 2026-07-17T20:18:03.731422+00:00
Summary
Audio-FLAN is a large-scale instruction-tuning dataset for unified audio understanding and generation, integrating 50+ source datasets across speech, music, and general audio domains. It covers tasks like speech recognition, text-to-speech, acoustic scene classification, audio event recognition, and music generation, organized with structured metadata (instruction, input, output, task type) for training audio-language models.
Keywords
audio speech music sound instruction-tuning text-to-speech asr audio-generation
Topics
Audio / Speech / Music
Research notes
- Requires sharing contact information for access. Audio included only for datasets with permissive licenses (CC-BY, Apache-2.0, MIT); metadata-only for restrictive licenses. Paper: arxiv 2502.16584.