← Back to explorer

Datasets Catalog Spreadsheet

Type
collection
Venue
Google Sheets (shared publicly)
Year
2026
Source
google-docs
Access
free
Language
English
Added
2026-07-17T20:18:03.699032+00:00
Verified
2026-07-17T20:18:03.699032+00:00

Summary

This is a publicly shared Google Sheets spreadsheet that serves as a comprehensive catalog of AI/ML training datasets, with detailed metadata columns including domain, purpose, language distribution, author distribution, scale, bias distribution, text quality, license, cost, collection method, sample size distribution, long context dependency, and processing status. It includes entries for datasets such as Institutional Books (242B tokens), British Library Books (25M pages), Sci-Base (600B tokens), OpenR1-Math-220k, DeepScaleR, Harvard Library Public Domain Corpus, Public Domain Poetry, Mixture of Thoughts, Open Math Reasoning, OLMo Mix, Essential-Web 1.0, WildChat, and others, making it a valuable reference for dataset selection in LLM training.

Keywords

catalog datasets metadata reference llm-training spreadsheet data-selection

Topics

AI/ML / Data Science

Research notes

  • Google Sheets URL requires JavaScript to render; CSV export was used to read content. The spreadsheet includes datasets like Institutional Books, British Library Books, Sci-Base, OpenR1-Math-220k, DeepScaleR, Harvard Library Public Domain Corpus, Public Domain Poetry, AMC Problems, Mixture of Thoughts, Eurlex Multilingual, Open Math Reasoning, OLMo Mix, Essential-Web 1.0, and WildChat. The original compiler/author is not identified in the sheet itself.