← Back to explorer

eurlex-multilingual

Type
dataset
Venue
HuggingFace (mteb)
Year
2026
Source
huggingface
Access
free
Language
23 EU languages (Bulgarian, Czech, Danish, German, Greek, English, Spanish, Estonian, Finnish, French, Croatian, Hungarian, Italian, Lithuanian, Latvian, Maltese, Dutch, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish)
Added
2026-07-17T20:18:03.619678+00:00
Verified
2026-07-17T20:18:03.619678+00:00

Summary

eurlex-multilingual is an MTEB benchmark dataset containing European Union legal documents (regulations, decisions, directives) in 23 EU languages, each annotated with multi-label EUROVOC concept labels (21 concepts). The dataset has 23 language subsets (ranging from ~15k to ~65k rows each) with train/validation/test splits, and is used to evaluate multilingual text embedding and classification models. It is based on the EURLEX dataset and associated with arXiv papers 2109.00904, 2502.13595, and 2210.07316.

Keywords

eurlex multilingual classification legal eu-law eurovoc mteb embeddings benchmark

Topics

NLP / Classification

Research notes

  • Part of the MTEB benchmark suite. The dataset is derived from EURLEX, the official legal text database of the EU.