← Back to explorer

HeQ - Hebrew Question Answering Dataset

Type
dataset
Venue
NNLP-IL (National NLP Initiative Israel)
Year
2026
Source
github
Access
free
Language
Hebrew
Added
2026-07-17T20:18:03.584523+00:00
Verified
2026-07-17T20:18:03.584523+00:00

Summary

HeQ is a question answering dataset in Modern Hebrew consisting of 30,147 questions, following the format and crowdsourcing methodology of SQuAD and the earlier ParaShoot dataset. Crowdworkers formulated and answered reading comprehension questions based on paragraphs sourced from Hebrew Wikipedia and Geektime (an Israeli technology news site), with answers being text spans extracted from the relevant paragraphs. The dataset is designed to address the challenges of extractive QA in Hebrew, a morphologically rich language (MRL) where word boundaries may not align with semantic units, making QA more challenging than in morphologically simpler languages like English.

Keywords

hebrew question-answering squad reading-comprehension mrl nlp extractive-qa crowdsourced

Topics

NLP / Question Answering / Hebrew

Research notes

  • GitHub repo is accessible. Also available on HuggingFace (amirdnc/HeQ, Etelis/HeQ_v1). Published at Findings of EMNLP 2023 as 'HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark'. The annotation tool code is in a separate repo (NNLP-IL/Parashoot-Tagging).