← Back to explorer

Ubuntu Dialogue Corpus

Type
corpus
Venue
McGill University / Kaggle (rtatman mirror)
Year
2026
Source
kaggle
Access
free
Language
English
Added
2026-07-17T20:18:03.594565+00:00
Verified
2026-07-17T20:18:03.594565+00:00

Summary

The Ubuntu Dialogue Corpus contains almost 1 million two-person multi-turn dialogues extracted from Ubuntu chat logs used for technical support, totaling over 7 million utterances and 100 million words. Introduced by Lowe et al. at SIGDial 2015, it was designed to bridge the gap in large-scale datasets for training neural network-based dialogue managers and includes benchmark tasks for next-response selection.

Keywords

dialogue multi-turn ubuntu chat-logs conversational response-selection nlp benchmark

Topics

NLP / Dialogue Systems

Research notes

  • Kaggle page was JS-blocked (reCAPTCHA); info gathered from original McGill dataset site and SIGDial 2015 paper. v2.0 (recommended) available via rkadlec/ubuntu-ranking-dataset-creator. Also mirrored on HuggingFace as ubuntu-dialogs-corpus/ubuntu_dialogs_corpus.