← Back to explorer

Harvard Library Public Domain Corpus

Type
corpus
Venue
Harvard Library / Institutional Data Initiative
Year
2026
Source
harvard
Access
restricted
Language
254 languages (primarily English, Czech, Icelandic, Welsh, and others)
Added
2026-07-17T20:18:03.600607+00:00
Verified
2026-07-17T20:18:03.600607+00:00

Summary

The Harvard Library Public Domain Corpus is a collection of approximately 1 million digitized public domain books (983,004 volumes, ~242B tokens) originally scanned through Harvard Library's participation in the Google Books project beginning in 2006. The corpus spans 254 languages and 386M pages of text (available in original and post-processed OCR formats), and was refined through collection-level deduplication and OCR post-processing by the Institutional Data Initiative, making it one of the largest lawful book corpora available for LLM training.

Keywords

public-domain books harvard google-books ocr multilingual llm-training historical institutional

Topics

Public Domain Books / Literature / History

Research notes

  • Access currently restricted to nonprofit/educational/research uses via request. Also released as 'Institutional Books 1.0' on HuggingFace (institutional/institutional-books-1.0-metadata). Funded by Microsoft and OpenAI. Described in arXiv:2506.08300.