Hebrew Projectbenyehuda
- Type
- dataset
- Venue
- projectbenyehuda
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- he
- Added
- 2026-07-17T19:58:59.399957+00:00
- Verified
- 2026-07-17T19:58:59.399957+00:00
Summary
This repository contains a dump of thousands of public domain works in Hebrew, from Project Ben-Yehuda, in plaintext UTF-8 files, with and without diacritics (nikkud), and in HTML files. The pseudocatalogue.csv file is a list of titles, authors, genres, and file paths, to help you process the dump.
Keywords
hf-dataset language-modeling masked-language-modeling expert-generated found monolingual original
Topics
Language Modeling
Research notes
- downloads=145; likes=4