SYNTH - generalist open data and environment
- Type
- dataset
- Venue
- PleIAs
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en, fr, it, es, de, pl, nl, la
- Added
- 2026-07-17T19:59:07.997161+00:00
- Verified
- 2026-07-17T19:59:07.997161+00:00
Summary
SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the *Structured Wikipedia* dataset from Wikimedia Enterprise.
Keywords
hf-dataset language-modeling · -nlp---zero-shot · -summarization · -math parquet text datasets dask polars mlcroissant wikipedia art math writing
Topics
Language Modeling, NLP / Zero-shot, Summarization, Math
Research notes
- downloads=6477; likes=272