Cornell Movie-Dialogs Corpus
- Type
- corpus
- Venue
- Cornell University / Kaggle (rajathmc)
- Year
- 2026
- Source
- kaggle
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.680633+00:00
- Verified
- 2026-07-17T20:18:03.680633+00:00
Summary
The Cornell Movie-Dialogs Corpus contains 220,579 conversational exchanges between 10,292 pairs of movie characters across 617 movies, totaling 304,713 utterances. It includes rich metadata such as movie genres, release years, IMDB ratings, character gender, and credit positions, making it a standard benchmark for open-domain dialogue generation and conversational style coordination. It was distributed alongside the paper 'Chameleons in Imagined Conversations' (ACL 2011).
Keywords
dialogue movies conversational-ai nlp fiction metadata-rich
Topics
NLP / Dialogue
Research notes
- Kaggle mirror by Rajath Chidananda; original hosted at Cornell CS. The Kaggle page is reCAPTCHA-blocked but the dataset is well-documented at the original Cornell source and via ConvoKit.