Internet Archive Dataset
- Type
- dataset
- Venue
- Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English (primarily)
- Added
- 2026-07-17T20:18:03.664275+00:00
- Verified
- 2026-07-17T20:18:03.664275+00:00
Summary
A dataset published by Nick Saga (nick007x) that contains content harvested from the Internet Archive. Based on the creator's pattern of scraping and archiving large-scale web data (GitHub, ArXiv, Reddit, Hacker News), this dataset likely contains OCR-processed texts, books, or documents from the Internet Archive's digital library, intended for use in language model pretraining or text-based research.
Keywords
internet-archive books public-domain ocr web-crawl archive
Topics
Public Domain Books / Web
Research notes
- The HuggingFace page returned HTTP 401 (gated or restricted access). The dataset is not listed on the creator's public profile page, suggesting it may have been renamed, made private, or deleted. Information inferred from the dataset name and the creator's known pattern of archiving web-scale data.