HeQ - Hebrew Question Answering Dataset
- Type
- dataset
- Venue
- NNLP-IL (National NLP Initiative Israel)
- Year
- 2026
- Source
- github
- Access
- free
- Language
- Hebrew
- Added
- 2026-07-17T20:18:03.584523+00:00
- Verified
- 2026-07-17T20:18:03.584523+00:00
Summary
HeQ is a question answering dataset in Modern Hebrew consisting of 30,147 questions, following the format and crowdsourcing methodology of SQuAD and the earlier ParaShoot dataset. Crowdworkers formulated and answered reading comprehension questions based on paragraphs sourced from Hebrew Wikipedia and Geektime (an Israeli technology news site), with answers being text spans extracted from the relevant paragraphs. The dataset is designed to address the challenges of extractive QA in Hebrew, a morphologically rich language (MRL) where word boundaries may not align with semantic units, making QA more challenging than in morphologically simpler languages like English.
Keywords
hebrew question-answering squad reading-comprehension mrl nlp extractive-qa crowdsourced
Topics
NLP / Question Answering / Hebrew
Research notes
- GitHub repo is accessible. Also available on HuggingFace (amirdnc/HeQ, Etelis/HeQ_v1). Published at Findings of EMNLP 2023 as 'HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark'. The annotation tool code is in a separate repo (NNLP-IL/Parashoot-Tagging).