Datasets Catalog Spreadsheet
- Type
- collection
- Venue
- Google Sheets (shared publicly)
- Year
- 2026
- Source
- google-docs
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.699032+00:00
- Verified
- 2026-07-17T20:18:03.699032+00:00
Summary
This is a publicly shared Google Sheets spreadsheet that serves as a comprehensive catalog of AI/ML training datasets, with detailed metadata columns including domain, purpose, language distribution, author distribution, scale, bias distribution, text quality, license, cost, collection method, sample size distribution, long context dependency, and processing status. It includes entries for datasets such as Institutional Books (242B tokens), British Library Books (25M pages), Sci-Base (600B tokens), OpenR1-Math-220k, DeepScaleR, Harvard Library Public Domain Corpus, Public Domain Poetry, Mixture of Thoughts, Open Math Reasoning, OLMo Mix, Essential-Web 1.0, WildChat, and others, making it a valuable reference for dataset selection in LLM training.
Keywords
catalog datasets metadata reference llm-training spreadsheet data-selection
Topics
AI/ML / Data Science
Research notes
- Google Sheets URL requires JavaScript to render; CSV export was used to read content. The spreadsheet includes datasets like Institutional Books, British Library Books, Sci-Base, OpenR1-Math-220k, DeepScaleR, Harvard Library Public Domain Corpus, Public Domain Poetry, AMC Problems, Mixture of Thoughts, Eurlex Multilingual, Open Math Reasoning, OLMo Mix, Essential-Web 1.0, and WildChat. The original compiler/author is not identified in the sheet itself.