Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)
- Type
- paper
- Venue
- arXiv / UFMG
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:50:00Z
- Verified
- 2026-08-14T20:50:00Z
Summary
Largest claimed public Discord dataset: 2,052,206,308 messages from 4,735,057 users across 3,167 public Discovery servers (~10% of 31,673 servers listed as of 2024-11-17), spanning 2015-05-13 to 2024-12-17. Collected via Discord API; usernames pseudonymized (mimesis), IDs SHA-256 truncated to 12 chars. Per-server JSON plus servers_metadata. 17% of messages from bots. English-US dominates preferred_locale (1705 servers) with Spanish, French, Portuguese also present. Gaming still the top description keyword (~15%). Authors argue Discord is understudied vs Twitter/Reddit and useful for user-driven moderation research. Paper submitted to ICWSM 2025.
Keywords
discord · social-media · public-chat · moderation · bots · icwsm · zenodo · ufmg
Topics
computational social science, Discord, public chat dumps
Research notes
- Primary: arxiv abs (cs.SI/cs.DB; Submitted to ICWSM 2025). Paper cites Zenodo DOI 10.5281/zenodo.14658505; that record 404 at 2026-08-14 check. HF paper page links unofficial mirrors fvdfs41/Discord-Unveiled and SaisExperiments/Discord-Unveiled-Compressed — not treated as official. Anonymized public Discovery servers only. Dataset not in datasets_local.csv. Discord posted abs.