← Back to explorer

Infinity-Doc2-5M

Type
dataset
Venue
infly / Hugging Face
Year
2026
Source
huggingface
Access
free
Language
English, Chinese
Added
2026-07-17T20:18:35.160169+00:00
Verified
2026-07-17T20:18:35.160169+00:00

Summary

A large-scale, high-quality training dataset of 5 million document pages specifically designed for document parsing tasks, covering academic papers, textbooks, exam papers, magazines, newspapers, financial reports, and other real-world document types in both Chinese and English. It provides multi-level annotations from block-level to page-level, including element bounding boxes, categories (titles, text, tables, formulas, headers, footers), content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. The dataset was constructed using a scalable synthesis engine with a controllable rendering framework and iterative refinement loop, and is used to train the Infinity-Parser2-Pro and Infinity-Parser2-Flash models.

Keywords

document-parsing ocr pdf layout-analysis multimodal bilingual chinese english table-recognition formula-recognition

Topics

Document AI / OCR / Multimodal

Research notes

  • Paper: arxiv 2607.07836 'Infinity-Parser2 Technical Report'. The HF demo subset shows 130 rows but the full dataset contains 5M samples. Supports 8 co-trained objectives: document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding.