Wikipedia Structured Contents (Kaggle)
- Type
- corpus
- Venue
- Kaggle (Wikimedia Foundation)
- Year
- 2026
- Source
- kaggle
- Access
- free
- Language
- English, French
- Added
- 2026-07-17T20:18:03.621377+00:00
- Verified
- 2026-07-17T20:18:03.621377+00:00
Summary
The Wikipedia Structured Contents dataset on Kaggle is a beta release by Wikimedia Enterprise featuring structured, pre-parsed Wikipedia article content in English and French, formatted as clean JSON/Parquet for machine learning use. Instead of scraping raw wikitext, users get developer-friendly representations with abstracts, short descriptions, infobox-style key-value data, image links, and segmented article sections. It is powered by the Snapshot API's Structured Contents beta and contains ~7.6M English articles (34.6 GiB) and ~2.9M French articles (9.8 GiB). The same data is also available on HuggingFace as wikimedia/structured-wikipedia.
Keywords
wikipedia structured-data parquet knowledge multilingual wikimedia pretraining nlp
Topics
NLP / Knowledge
Research notes
- The Kaggle page was previously blocked by reCAPTCHA for automated research. The dataset was updated in May 2026 with new features from the Structured Contents Initiative. Also mirrored on HuggingFace as wikimedia/structured-wikipedia.