daVinci-LLM Data
- Type
- corpus
- Venue
- SII-GAIR (Shanghai Innovation Institute - Global AI Research in NLP)
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English
- Added
- 2026-07-17T20:18:03.772532+00:00
- Verified
- 2026-07-17T20:18:03.772532+00:00
Summary
daVinci-LLM Data is a subset of the daVinci-LLM training corpus released under the 'Data Darwinism' taxonomy, which classifies data by processing depth. It includes classified web corpus (L3, 4.28T from Nemotron-CC-v1), refined math corpora (L4, from MegaMath), and QA datasets (L5, synthetic reasoning data in math and science). The release aims to make data curation decisions explicit and transparent, with each source annotated with a Darwin Level reflecting how deeply it has been processed.
Keywords
pretraining web-data math code science data-darwinism curation corpus
Topics
General / Math / Science / Code
Research notes
- Requires agreeing to share contact information to access. Associated with arXiv:2603.27164 and arXiv:2602.07824. Code portion planned for future release.