Webis Gmane Email Corpus 2019
- Type
- corpus
- Venue
- Webis / Bauhaus-Universität Weimar
- Year
- 2026
- Source
- webis
- Access
- restricted
- Language
- Multiple (language detected per email)
- Added
- 2026-07-17T20:18:03.684098+00:00
- Verified
- 2026-07-17T20:18:03.684098+00:00
Summary
The Webis Gmane Email Corpus 2019 is a dataset of over 153 million parsed and segmented emails crawled from gmane.io between February and May 2019, covering more than 20 years of public mailing list discussions across 14,699 lists. Each email is segmented into semantically consistent components (paragraphs, quotations, signatures, salutations, etc.) using the Chipmunk neural segmentation model with 96% accuracy across 15 segment classes. Published as a resource at ACL 2020.
Keywords
email mailing-lists dialog-analysis segmentation nlp large-scale
Topics
NLP / Email / Dialogue
Research notes
- Available on Zenodo (doi:10.5281/zenodo.3766984). Files are restricted to researchers. Data is in Elasticsearch bulk JSON format with Gzip compression. Related code at github.com/webis-de/acl20-crawling-mailing-lists.