Viet-Handwriting-OCR-v2
- Type
- dataset
- Venue
- 5CD-AI (Fifth Civil Defender) / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- Vietnamese
- Added
- 2026-07-17T20:18:35.176443+00:00
- Verified
- 2026-07-17T20:18:35.176443+00:00
Summary
A dataset of 60,247 Vietnamese handwritten text images collected and curated for handwritten text recognition research, with images crawled from public internet sources and manually annotated by human labelers to ensure high diversity and transcription accuracy. The v2 release adds 35,844 new training samples compared to v1, significantly expanding the diversity and coverage of Vietnamese handwriting styles. All images are cropped to individual lines or single sentences with no personally identifiable information, and the dataset is released under a non-commercial license for academic purposes and Vietnamese OCR/AI technology research.
Keywords
ocr text-recognition handwriting handwriting-recognition vietnamese low-resource-language htr
Topics
Vision / OCR / Vietnamese NLP
Research notes
- Requires agreeing to share contact information to access. Paper: arxiv 2408.12480 'Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese'. Compliant with Vietnam's Decree 13/2023/ND-CP on Personal Data Protection. Takedown policy: contact huynhgiabaoa2@gmail.com or dtkhangbk@gmail.com for removal within 24-48 hours. Actively scaling to hundreds of thousands of human-verified samples.