ArXiv Papers
- Type
- dataset
- Venue
- Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English (primarily)
- Added
- 2026-07-17T20:18:03.661266+00:00
- Verified
- 2026-07-17T20:18:03.661266+00:00
Summary
A massive scientific corpus containing 2.55 million arXiv papers with complete metadata across all academic domains, totaling approximately 4.6 TB. Each record includes the arXiv ID, title, authors, submission date, comments, primary subject category, subject classifications, DOI, abstract, and file path to the source PDF/LaTeX. It is designed for training models on academic reasoning, literature review, and scientific knowledge mining.
Keywords
arxiv scientific-papers academic research large-scale metadata
Topics
Science / NLP
Research notes
- Published by Nick Saga (nick007x) on HuggingFace. Includes abstracts and full metadata for all arXiv domains.