← Back to explorer

Usenet Corpus 1980–2013

Type
dataset
Venue
OwnedByDanes
Year
2026
Source
huggingface
Access
free
Language
en, pl, nl, es, fr, it, de, ru, ja, pt
Added
2026-07-17T20:02:11.229071+00:00
Verified
2026-07-17T20:02:11.229071+00:00

Summary

This dataset contains **408 million cleaned and deduplicated Usenet posts** spanning 1980–2013 across 18,347 newsgroups. It is sourced from one of the largest privately held Usenet corpora and has been rigorously processed for modern AI training use cases.

Keywords

hf-dataset language-modeling · -chat---dialogue json text datasets dask polars mlcroissant usenet internet-history forums long-form-text pre-web conversational historical pretraining

Topics

Language Modeling, Chat / Dialogue

Research notes

  • downloads=119; likes=20