Discord-Data
- Type
- dataset
- Venue
- Kaggle / JEF1056
- Year
- 2026
- Source
- kaggle
- Access
- free
- Language
- English (primarily)
- Added
- 2026-07-17T20:18:03.657912+00:00
- Verified
- 2026-07-17T20:18:03.657912+00:00
Summary
Discord-Data is a large corpus of chat messages scraped from many public Discord servers, containing on the order of 110-300 million messages (the creator references 'over 300 million messages' while a downstream analysis cites ~110 million). It is intended for NLP/conversation research and is accompanied by a companion cleaning repository (JEF1056/clean-discord) that performs detoxification, conversation-turn splitting, and dataset-slice generation. The dataset has been used for chat-log statistical and linguistic analysis.
Keywords
discord chat conversations nlp social-media scraped kaggle corpus
Topics
NLP / Conversational / Social Media
Research notes
- The Kaggle page crashed during automated fetch (reCAPTCHA/JS). Info synthesized from the creator's GitHub repo (JEF1056/clean-discord) and a downstream analysis repo (b7leung/Chat-Log-Statistical-Linguistic-Analysis). Message-count estimates vary between sources (110M vs 300M). License not explicitly stated; content is scraped from public Discord servers.