Institutional Books 1.0
- Type
- corpus
- Venue
- Institutional Data Initiative (Harvard University)
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- 254 unique volume-level languages (primarily English, with many others)
- Added
- 2026-07-17T20:18:03.626990+00:00
- Verified
- 2026-07-17T20:18:03.626990+00:00
Summary
Institutional Books 1.0 is a corpus of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative (IDI). It contains 242 billion o200k_base tokens across 386 million pages in 254 languages, with extensive volume-level metadata including OCR quality, language distributions, text analysis, and HathiTrust data extensions. The books were published largely in the 19th and 20th centuries. Access is restricted to noncommercial use with no redistribution allowed, requiring agreement to IDI's Early-Access Terms of Use.
Keywords
books public-domain harvard google-books pretraining ocr multilingual corpus
Topics
Public Domain Books
Research notes
- Gated dataset requiring contact information agreement. Commercial use requires contacting contact@institutional.org. Attribution must include 'Institutional Books provided by the Institutional Data Initiative with source material contributed by Harvard Library.' Associated with arXiv 2506.08300.