← Back to explorer

Wikipedia Structured Contents (Kaggle)

Type
corpus
Venue
Kaggle (Wikimedia Foundation)
Year
2026
Source
kaggle
Access
free
Language
English, French
Added
2026-07-17T20:18:03.621377+00:00
Verified
2026-07-17T20:18:03.621377+00:00

Summary

The Wikipedia Structured Contents dataset on Kaggle is a beta release by Wikimedia Enterprise featuring structured, pre-parsed Wikipedia article content in English and French, formatted as clean JSON/Parquet for machine learning use. Instead of scraping raw wikitext, users get developer-friendly representations with abstracts, short descriptions, infobox-style key-value data, image links, and segmented article sections. It is powered by the Snapshot API's Structured Contents beta and contains ~7.6M English articles (34.6 GiB) and ~2.9M French articles (9.8 GiB). The same data is also available on HuggingFace as wikimedia/structured-wikipedia.

Keywords

wikipedia structured-data parquet knowledge multilingual wikimedia pretraining nlp

Topics

NLP / Knowledge

Research notes

  • The Kaggle page was previously blocked by reCAPTCHA for automated research. The dataset was updated in May 2026 with new features from the Structured Contents Initiative. Also mirrored on HuggingFace as wikimedia/structured-wikipedia.