← Back to explorer

Viet-Handwriting-OCR-v2

Type
dataset
Venue
5CD-AI (Fifth Civil Defender) / Hugging Face
Year
2026
Source
huggingface
Access
restricted
Language
Vietnamese
Added
2026-07-17T20:18:35.176443+00:00
Verified
2026-07-17T20:18:35.176443+00:00

Summary

A dataset of 60,247 Vietnamese handwritten text images collected and curated for handwritten text recognition research, with images crawled from public internet sources and manually annotated by human labelers to ensure high diversity and transcription accuracy. The v2 release adds 35,844 new training samples compared to v1, significantly expanding the diversity and coverage of Vietnamese handwriting styles. All images are cropped to individual lines or single sentences with no personally identifiable information, and the dataset is released under a non-commercial license for academic purposes and Vietnamese OCR/AI technology research.

Keywords

ocr text-recognition handwriting handwriting-recognition vietnamese low-resource-language htr

Topics

Vision / OCR / Vietnamese NLP

Research notes

  • Requires agreeing to share contact information to access. Paper: arxiv 2408.12480 'Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese'. Compliant with Vietnam's Decree 13/2023/ND-CP on Personal Data Protection. Takedown policy: contact huynhgiabaoa2@gmail.com or dtkhangbk@gmail.com for removal within 24-48 hours. Actively scaling to hundreds of thousands of human-verified samples.