Infinity-Doc2-5M
- Type
- dataset
- Venue
- infly / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English, Chinese
- Added
- 2026-07-17T20:18:35.160169+00:00
- Verified
- 2026-07-17T20:18:35.160169+00:00
Summary
A large-scale, high-quality training dataset of 5 million document pages specifically designed for document parsing tasks, covering academic papers, textbooks, exam papers, magazines, newspapers, financial reports, and other real-world document types in both Chinese and English. It provides multi-level annotations from block-level to page-level, including element bounding boxes, categories (titles, text, tables, formulas, headers, footers), content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. The dataset was constructed using a scalable synthesis engine with a controllable rendering framework and iterative refinement loop, and is used to train the Infinity-Parser2-Pro and Infinity-Parser2-Flash models.
Keywords
document-parsing ocr pdf layout-analysis multimodal bilingual chinese english table-recognition formula-recognition
Topics
Document AI / OCR / Multimodal
Research notes
- Paper: arxiv 2607.07836 'Infinity-Parser2 Technical Report'. The HF demo subset shows 130 rows but the full dataset contains 5M samples. Supports 8 co-trained objectives: document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding.