← Back to explorer

Discord-Data (jef1056)

Type
corpus
Venue
Kaggle
Year
2026
Source
kaggle
Access
free
Language
English
Added
2026-07-17T20:18:03.596496+00:00
Verified
2026-07-17T20:18:03.596496+00:00

Summary

Discord-Data is a large-scale dataset of over 300 million chat messages scraped from public Discord servers, compiled by JEF1056 (JFan) and hosted on Kaggle. It was designed for NLP research on informal online conversation and includes a companion cleaning repository (JEF1056/clean-discord) for filtering, detoxing, and creating conversation splits, taking approximately 6 hours to process on a 4-core machine.

Keywords

discord chat social-media nlp dialogue scraped conversational large-scale

Topics

NLP / Chat / Social Media

Research notes

  • Kaggle page was JS-blocked; info gathered from JEF1056/clean-discord GitHub repo and b7leung/Chat-Log-Statistical-Linguistic-Analysis. Dataset size varies by source (300M raw vs ~110M cited in one analysis).