FinePDFs
- Type
- corpus
- Venue
- HuggingFace (HuggingFaceFW / FineData)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- 1733 languages (including English, Arabic, French, German, Spanish, Chinese, and many low-resource languages)
- Added
- 2026-07-17T20:18:03.634840+00:00
- Verified
- 2026-07-17T20:18:03.634840+00:00
Summary
FinePDFs is the largest publicly available LLM pretraining corpus sourced exclusively from PDF documents, containing approximately 3 trillion tokens across 475 million documents in 1733 languages. The data was sourced from 105 CommonCrawl snapshots spanning summer 2013 to February 2025, refetched from the internet, and processed using the datatrove library with careful deduplication, OCR, and filtering. The dataset (3.65 TB) includes a companion FinePDFs-EDU variant with educational content filtering. When mixed with FW-Edu+DCLM, FinePDFs outperforms even Nemotron-CC v2 in pretraining ablations. The processing pipeline and code are available on GitHub at huggingface/finepdfs.
Keywords
pdf pretraining corpus commoncrawl multilingual ocr fineweb llm datatrove
Topics
NLP / Pretraining Data
Research notes
- Associated with arXiv papers 2506.18421 and 2109.07445. Includes PII filtering and copyright opt-out mechanisms. The English subset alone has 207M rows. A companion FinePDFs-EDU dataset applies additional educational quality filtering, removing ~96% of the initial data.