← Back to explorer

Institutional Books 1.0

Type
corpus
Venue
Institutional Data Initiative (Harvard University)
Year
2026
Source
huggingface
Access
restricted
Language
254 unique volume-level languages (primarily English, with many others)
Added
2026-07-17T20:18:03.626990+00:00
Verified
2026-07-17T20:18:03.626990+00:00

Summary

Institutional Books 1.0 is a corpus of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative (IDI). It contains 242 billion o200k_base tokens across 386 million pages in 254 languages, with extensive volume-level metadata including OCR quality, language distributions, text analysis, and HathiTrust data extensions. The books were published largely in the 19th and 20th centuries. Access is restricted to noncommercial use with no redistribution allowed, requiring agreement to IDI's Early-Access Terms of Use.

Keywords

books public-domain harvard google-books pretraining ocr multilingual corpus

Topics

Public Domain Books

Research notes

  • Gated dataset requiring contact information agreement. Commercial use requires contacting contact@institutional.org. Attribution must include 'Institutional Books provided by the Institutional Data Initiative with source material contributed by Harvard Library.' Associated with arXiv 2506.08300.