← Back to explorer

Discord-Data

Type
dataset
Venue
Kaggle / JEF1056
Year
2026
Source
kaggle
Access
free
Language
English (primarily)
Added
2026-07-17T20:18:03.657912+00:00
Verified
2026-07-17T20:18:03.657912+00:00

Summary

Discord-Data is a large corpus of chat messages scraped from many public Discord servers, containing on the order of 110-300 million messages (the creator references 'over 300 million messages' while a downstream analysis cites ~110 million). It is intended for NLP/conversation research and is accompanied by a companion cleaning repository (JEF1056/clean-discord) that performs detoxification, conversation-turn splitting, and dataset-slice generation. The dataset has been used for chat-log statistical and linguistic analysis.

Keywords

discord chat conversations nlp social-media scraped kaggle corpus

Topics

NLP / Conversational / Social Media

Research notes

  • The Kaggle page crashed during automated fetch (reCAPTCHA/JS). Info synthesized from the creator's GitHub repo (JEF1056/clean-discord) and a downstream analysis repo (b7leung/Chat-Log-Statistical-Linguistic-Analysis). Message-count estimates vary between sources (110M vs 300M). License not explicitly stated; content is scraped from public Discord servers.