OCR Synthetic Multilingual v1
- Type
- dataset
- Venue
- nvidia
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en, ja, ko, ru, zh
- Added
- 2026-07-17T20:02:03.856047+00:00
- Verified
- 2026-07-17T20:02:03.856047+00:00
Summary
Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of [SynthDoG](https://github.com/clovaai/donut/tree/master/synthdog) (Synthetic Document Generator), originally introduced in the [Donut](https://github.com/clovaai/donut) project by Kim et al.
Keywords
hf-dataset vision---detection ocr text-detection text-recognition synthetic-data synthdog hdf5 nvidia nemotron
Topics
Vision / Detection
Research notes
- downloads=18714; likes=50