← Back to explorer

common_corpus

Type
dataset
Venue
PleIAs
Year
2026
Source
huggingface
Access
free
Language
en, fr, de, zh, it, es, ja, pl, la, nl, ru, ar, ko
Added
2026-07-17T19:56:42.048139+00:00
Verified
2026-07-17T19:56:42.048139+00:00

Summary

Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.

Keywords

hf-dataset parquet tabular text datasets pandas polars mlcroissant has-paper

Research notes

  • downloads=81783; likes=409