Discord-Data (jef1056)
- Type
- corpus
- Venue
- Kaggle
- Year
- 2026
- Source
- kaggle
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.596496+00:00
- Verified
- 2026-07-17T20:18:03.596496+00:00
Summary
Discord-Data is a large-scale dataset of over 300 million chat messages scraped from public Discord servers, compiled by JEF1056 (JFan) and hosted on Kaggle. It was designed for NLP research on informal online conversation and includes a companion cleaning repository (JEF1056/clean-discord) for filtering, detoxing, and creating conversation splits, taking approximately 6 hours to process on a 4-core machine.
Keywords
discord chat social-media nlp dialogue scraped conversational large-scale
Topics
NLP / Chat / Social Media
Research notes
- Kaggle page was JS-blocked; info gathered from JEF1056/clean-discord GitHub repo and b7leung/Chat-Log-Statistical-Linguistic-Analysis. Dataset size varies by source (300M raw vs ~110M cited in one analysis).