common_corpus
- Type
- dataset
- Venue
- PleIAs
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en, fr, de, zh, it, es, ja, pl, la, nl, ru, ar, ko
- Added
- 2026-07-17T19:56:42.048139+00:00
- Verified
- 2026-07-17T19:56:42.048139+00:00
Summary
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Keywords
hf-dataset parquet tabular text datasets pandas polars mlcroissant has-paper
Research notes
- downloads=81783; likes=409