Dolma
- Type
- dataset
- Venue
- allenai
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T19:57:11.136439+00:00
- Verified
- 2026-07-17T19:57:11.136439+00:00
Summary
Dolma is a dataset of 3 trillion tokens from a diverse mix of web content, academic publications, code, books, and encyclopedic materials.
Keywords
hf-dataset language-modeling casual-lm llm has-paper
Topics
Language Modeling
Research notes
- downloads=3753; likes=1057