Ubuntu Dialogue Corpus
- Type
- corpus
- Venue
- McGill University / Kaggle (rtatman mirror)
- Year
- 2026
- Source
- kaggle
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.594565+00:00
- Verified
- 2026-07-17T20:18:03.594565+00:00
Summary
The Ubuntu Dialogue Corpus contains almost 1 million two-person multi-turn dialogues extracted from Ubuntu chat logs used for technical support, totaling over 7 million utterances and 100 million words. Introduced by Lowe et al. at SIGDial 2015, it was designed to bridge the gap in large-scale datasets for training neural network-based dialogue managers and includes benchmark tasks for next-response selection.
Keywords
dialogue multi-turn ubuntu chat-logs conversational response-selection nlp benchmark
Topics
NLP / Dialogue Systems
Research notes
- Kaggle page was JS-blocked (reCAPTCHA); info gathered from original McGill dataset site and SIGDial 2015 paper. v2.0 (recommended) available via rkadlec/ubuntu-ranking-dataset-creator. Also mirrored on HuggingFace as ubuntu-dialogs-corpus/ubuntu_dialogs_corpus.