← Back to explorer

Internet Archive Dataset

Type
dataset
Venue
Hugging Face
Year
2026
Source
huggingface
Access
restricted
Language
English (primarily)
Added
2026-07-17T20:18:03.664275+00:00
Verified
2026-07-17T20:18:03.664275+00:00

Summary

A dataset published by Nick Saga (nick007x) that contains content harvested from the Internet Archive. Based on the creator's pattern of scraping and archiving large-scale web data (GitHub, ArXiv, Reddit, Hacker News), this dataset likely contains OCR-processed texts, books, or documents from the Internet Archive's digital library, intended for use in language model pretraining or text-based research.

Keywords

internet-archive books public-domain ocr web-crawl archive

Topics

Public Domain Books / Web

Research notes

  • The HuggingFace page returned HTTP 401 (gated or restricted access). The dataset is not listed on the creator's public profile page, suggesting it may have been renamed, made private, or deleted. Information inferred from the dataset name and the creator's known pattern of archiving web-scale data.