← Back to explorer

OPUS - Open Parallel Corpus Collection

Type
corpus
Venue
University of Helsinki / Uppsala University
Year
2026
Source
opus
Access
free
Language
1005 languages
Added
2026-07-17T20:18:03.668245+00:00
Verified
2026-07-17T20:18:03.668245+00:00

Summary

OPUS is the largest freely available collection of parallel (translated) text corpora on the web, containing 1,214 individual corpora with over 102.9 billion sentence pairs across 1,005 languages. Major sub-corpora include OpenSubtitles (27.2B sentence pairs), NLLB (22.7B), CCMatrix (17.1B), and ParaCrawl (4.6B). The collection is compiled from open-source documentation, movie subtitles, web-crawled data, and institutional translations, with automatic sentence alignment and linguistic annotation.

Keywords

parallel-corpus translation multilingual machine-translation bilingual open-source large-scale

Topics

NLP / Translation

Research notes

  • Created and maintained by Jörg Tiedemann. Originally at Uppsala University, now at University of Helsinki (Helsinki-NLP). Cite: Tiedemann, 2012, 'Parallel Data, Tools and Interfaces in OPUS' (LREC 2012). GitHub: Helsinki-NLP/OPUS.