← Back to explorer

SYNTH - generalist open data and environment

Type
dataset
Venue
PleIAs
Year
2026
Source
huggingface
Access
free
Language
en, fr, it, es, de, pl, nl, la
Added
2026-07-17T19:59:07.997161+00:00
Verified
2026-07-17T19:59:07.997161+00:00

Summary

SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the *Structured Wikipedia* dataset from Wikimedia Enterprise.

Keywords

hf-dataset language-modeling · -nlp---zero-shot · -summarization · -math parquet text datasets dask polars mlcroissant wikipedia art math writing

Topics

Language Modeling, NLP / Zero-shot, Summarization, Math

Research notes

  • downloads=6477; likes=272