OvisOCR2 Technical Report
- Type
- paper
- Venue
- arXiv / Alibaba ATH-MaaS
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:15:00Z
- Verified
- 2026-08-14T19:15:00Z
Summary
Alibaba ATH-MaaS compact document parser (cite lu2026ovisocr2): Qwen3.5-0.8B post-trained with a real+synthetic HTML-aligned data engine, then SFT, GRPO on a 4B teacher with text/formula/table rewards (edit distance, CDM, TEDS), on-policy distillation into 0.8B, and model soup. Given a page image, emits Markdown in reading order covering text, LaTeX, HTML tables, and visual-region bbox tags. OmniDocBench v1.6 overall 96.58 (first end-to-end to top a leaderboard long led by pipelines; PaddleOCR-VL-1.6 96.33). PureDocBench Avg3 75.06. In-house >1k pages overall 85.54; handwriting 72.28; complex-table missing rate 7.96% vs pipelines 13–17%. Apache-2.0 weights ATH-MaaS/OvisOCR2; inference via vLLM 0.22.1. Discord post is Zhidongxi news, not the authors.
Keywords
ovisocr2 · document-parsing · ocr · omnidocbench · qwen3.5 · alibaba · ath-maas · x
Topics
document parsing, OCR, end-to-end VLMs
Research notes
- Primary: arxiv abs 2607.13639. Discord/X https://x.com/Chinazhidx/status/2080578336716398782 via fxtwitter. Weights https://huggingface.co/ATH-MaaS/OvisOCR2 (Apache-2.0). Code https://github.com/ATH-MaaS/Ovis. Demo HF space ATH-MaaS/OvisOCR2. Training data not released. Open weights, not a new hosted corpus, so no datasets_local row.