← Back to explorer

ArXiv Papers

Type
dataset
Venue
Hugging Face
Year
2026
Source
huggingface
Access
free
Language
English (primarily)
Added
2026-07-17T20:18:03.661266+00:00
Verified
2026-07-17T20:18:03.661266+00:00

Summary

A massive scientific corpus containing 2.55 million arXiv papers with complete metadata across all academic domains, totaling approximately 4.6 TB. Each record includes the arXiv ID, title, authors, submission date, comments, primary subject category, subject classifications, DOI, abstract, and file path to the source PDF/LaTeX. It is designed for training models on academic reasoning, literature review, and scientific knowledge mining.

Keywords

arxiv scientific-papers academic research large-scale metadata

Topics

Science / NLP

Research notes

  • Published by Nick Saga (nick007x) on HuggingFace. Includes abstracts and full metadata for all arXiv domains.