Harvard Library Public Domain Corpus
- Type
- corpus
- Venue
- Harvard Library / Institutional Data Initiative
- Year
- 2026
- Source
- harvard
- Access
- restricted
- Language
- 254 languages (primarily English, Czech, Icelandic, Welsh, and others)
- Added
- 2026-07-17T20:18:03.600607+00:00
- Verified
- 2026-07-17T20:18:03.600607+00:00
Summary
The Harvard Library Public Domain Corpus is a collection of approximately 1 million digitized public domain books (983,004 volumes, ~242B tokens) originally scanned through Harvard Library's participation in the Google Books project beginning in 2006. The corpus spans 254 languages and 386M pages of text (available in original and post-processed OCR formats), and was refined through collection-level deduplication and OCR post-processing by the Institutional Data Initiative, making it one of the largest lawful book corpora available for LLM training.
Keywords
public-domain books harvard google-books ocr multilingual llm-training historical institutional
Topics
Public Domain Books / Literature / History
Research notes
- Access currently restricted to nonprofit/educational/research uses via request. Also released as 'Institutional Books 1.0' on HuggingFace (institutional/institutional-books-1.0-metadata). Funded by Microsoft and OpenAI. Described in arXiv:2506.08300.