← Back to explorer

FinePDFs

Type
corpus
Venue
HuggingFace (HuggingFaceFW / FineData)
Year
2026
Source
huggingface
Access
free
Language
1733 languages (including English, Arabic, French, German, Spanish, Chinese, and many low-resource languages)
Added
2026-07-17T20:18:03.634840+00:00
Verified
2026-07-17T20:18:03.634840+00:00

Summary

FinePDFs is the largest publicly available LLM pretraining corpus sourced exclusively from PDF documents, containing approximately 3 trillion tokens across 475 million documents in 1733 languages. The data was sourced from 105 CommonCrawl snapshots spanning summer 2013 to February 2025, refetched from the internet, and processed using the datatrove library with careful deduplication, OCR, and filtering. The dataset (3.65 TB) includes a companion FinePDFs-EDU variant with educational content filtering. When mixed with FW-Edu+DCLM, FinePDFs outperforms even Nemotron-CC v2 in pretraining ablations. The processing pipeline and code are available on GitHub at huggingface/finepdfs.

Keywords

pdf pretraining corpus commoncrawl multilingual ocr fineweb llm datatrove

Topics

NLP / Pretraining Data

Research notes

  • Associated with arXiv papers 2506.18421 and 2109.07445. Includes PII filtering and copyright opt-out mechanisms. The English subset alone has 207M rows. A companion FinePDFs-EDU dataset applies additional educational quality filtering, removing ~96% of the initial data.