eurlex-multilingual
- Type
- dataset
- Venue
- HuggingFace (mteb)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- 23 EU languages (Bulgarian, Czech, Danish, German, Greek, English, Spanish, Estonian, Finnish, French, Croatian, Hungarian, Italian, Lithuanian, Latvian, Maltese, Dutch, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish)
- Added
- 2026-07-17T20:18:03.619678+00:00
- Verified
- 2026-07-17T20:18:03.619678+00:00
Summary
eurlex-multilingual is an MTEB benchmark dataset containing European Union legal documents (regulations, decisions, directives) in 23 EU languages, each annotated with multi-label EUROVOC concept labels (21 concepts). The dataset has 23 language subsets (ranging from ~15k to ~65k rows each) with train/validation/test splits, and is used to evaluate multilingual text embedding and classification models. It is based on the EURLEX dataset and associated with arXiv papers 2109.00904, 2502.13595, and 2210.07316.
Keywords
eurlex multilingual classification legal eu-law eurovoc mteb embeddings benchmark
Topics
NLP / Classification
Research notes
- Part of the MTEB benchmark suite. The dataset is derived from EURLEX, the official legal text database of the EU.