OPUS - Open Parallel Corpus Collection
- Type
- corpus
- Venue
- University of Helsinki / Uppsala University
- Year
- 2026
- Source
- opus
- Access
- free
- Language
- 1005 languages
- Added
- 2026-07-17T20:18:03.668245+00:00
- Verified
- 2026-07-17T20:18:03.668245+00:00
Summary
OPUS is the largest freely available collection of parallel (translated) text corpora on the web, containing 1,214 individual corpora with over 102.9 billion sentence pairs across 1,005 languages. Major sub-corpora include OpenSubtitles (27.2B sentence pairs), NLLB (22.7B), CCMatrix (17.1B), and ParaCrawl (4.6B). The collection is compiled from open-source documentation, movie subtitles, web-crawled data, and institutional translations, with automatic sentence alignment and linguistic annotation.
Keywords
parallel-corpus translation multilingual machine-translation bilingual open-source large-scale
Topics
NLP / Translation
Research notes
- Created and maintained by Jörg Tiedemann. Originally at Uppsala University, now at University of Helsinki (Helsinki-NLP). Cite: Tiedemann, 2012, 'Parallel Data, Tools and Interfaces in OPUS' (LREC 2012). GitHub: Helsinki-NLP/OPUS.