← Back to explorer

daVinci-LLM Data

Type
corpus
Venue
SII-GAIR (Shanghai Innovation Institute - Global AI Research in NLP)
Year
2026
Source
huggingface
Access
restricted
Language
English
Added
2026-07-17T20:18:03.772532+00:00
Verified
2026-07-17T20:18:03.772532+00:00

Summary

daVinci-LLM Data is a subset of the daVinci-LLM training corpus released under the 'Data Darwinism' taxonomy, which classifies data by processing depth. It includes classified web corpus (L3, 4.28T from Nemotron-CC-v1), refined math corpora (L4, from MegaMath), and QA datasets (L5, synthetic reasoning data in math and science). The release aims to make data curation decisions explicit and transparent, with each source annotated with a Darwin Level reflecting how deeply it has been processed.

Keywords

pretraining web-data math code science data-darwinism curation corpus

Topics

General / Math / Science / Code

Research notes

  • Requires agreeing to share contact information to access. Associated with arXiv:2603.27164 and arXiv:2602.07824. Code portion planned for future release.