Knesset Corpus
- Type
- dataset
- Venue
- University of Haifa CLG / IAHLT / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- he
- Added
- 2026-08-14T16:32:00Z
- Verified
- 2026-08-14T16:32:00Z
Summary
Official Knesset plenary and committee protocols with sentence-level morphosyntax (POS, morphology, UD), named entities, and speaker/faction metadata. HF card: >35 million sentences, 1992–2024, size_categories 10M<n<100M, ~8.8 GB. LRE 2025 paper Table 1 on the digital subset: 42,063 protocols, 32.8M sentences, 384.6M tokens (plenary 1992–2022, committee 1998–2022). Released as JSONL.bz2/parquet plus CONLLU, CSV metadata, raw .doc/.pdf, and ParlaMint-IL TEI. CC BY-SA 4.0. Also on ElasticSearch/Kibana.
Keywords
hf-dataset hebrew knesset parliamentary ud ner jsonl politics diachronic gender
Topics
Hebrew NLP, parliamentary proceedings, computational social science
Research notes
- Primary: HF dataset card (cc-by-sa-4.0, he, 10M<n<100M, downloads=2790 likes=5 checked 2026-08-14) plus arXiv 2405.18115 / LRE paper. Discord #data-source-dump posted https://github.com/HaifaCLG/KnessetCorpus which redirects to the HF dataset. Paper counts (32.8M sentences / 384.6M tokens) describe the 1992–2022 digital subset; HF card is larger (35M+, through 2024). Gender split in the paper is ~81% male / 19% female sentences. Streaming load recommended. Also added to datasets_local.csv.