← Back to explorer

OCR Synthetic Multilingual v1

Type
dataset
Venue
nvidia
Year
2026
Source
huggingface
Access
free
Language
en, ja, ko, ru, zh
Added
2026-07-17T20:02:03.856047+00:00
Verified
2026-07-17T20:02:03.856047+00:00

Summary

Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of [SynthDoG](https://github.com/clovaai/donut/tree/master/synthdog) (Synthetic Document Generator), originally introduced in the [Donut](https://github.com/clovaai/donut) project by Kim et al.

Keywords

hf-dataset vision---detection ocr text-detection text-recognition synthetic-data synthdog hdf5 nvidia nemotron

Topics

Vision / Detection

Research notes

  • downloads=18714; likes=50