← Back to explorer

Webis Gmane Email Corpus 2019

Type
corpus
Venue
Webis / Bauhaus-Universität Weimar
Year
2026
Source
webis
Access
restricted
Language
Multiple (language detected per email)
Added
2026-07-17T20:18:03.684098+00:00
Verified
2026-07-17T20:18:03.684098+00:00

Summary

The Webis Gmane Email Corpus 2019 is a dataset of over 153 million parsed and segmented emails crawled from gmane.io between February and May 2019, covering more than 20 years of public mailing list discussions across 14,699 lists. Each email is segmented into semantically consistent components (paragraphs, quotations, signatures, salutations, etc.) using the Chipmunk neural segmentation model with 96% accuracy across 15 segment classes. Published as a resource at ACL 2020.

Keywords

email mailing-lists dialog-analysis segmentation nlp large-scale

Topics

NLP / Email / Dialogue

Research notes

  • Available on Zenodo (doi:10.5281/zenodo.3766984). Files are restricted to researchers. Data is in Elasticsearch bulk JSON format with Gzip compression. Related code at github.com/webis-de/acl20-crawling-mailing-lists.