← Back to explorer

Hebrew Projectbenyehuda

Type
dataset
Venue
projectbenyehuda
Year
2026
Source
huggingface
Access
free
Language
he
Added
2026-07-17T19:58:59.399957+00:00
Verified
2026-07-17T19:58:59.399957+00:00

Summary

This repository contains a dump of thousands of public domain works in Hebrew, from Project Ben-Yehuda, in plaintext UTF-8 files, with and without diacritics (nikkud), and in HTML files. The pseudocatalogue.csv file is a list of titles, authors, genres, and file paths, to help you process the dump.

Keywords

hf-dataset language-modeling masked-language-modeling expert-generated found monolingual original

Topics

Language Modeling

Research notes

  • downloads=145; likes=4