CoolFace
Datasetpublic

NHLOCAL/project-ben-yehuda

Project Ben-Yehuda Corpus Dataset Summary This dataset contains Hebrew literary texts sourced from Project Ben-Yehuda and structured for NLP, language modeling, text analysis, historical linguistics, and computational Hebrew research. Each record represents one text/work and includes the full text, a fixed source label, and structured metadata from the Project Ben-Yehuda catalogue. Current Version Dataset version: pby-2026.03Source snapshot: Project… See the full description on the dataset page: https://huggingface.co/datasets/NHLOCAL/project-ben-yehuda.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
2likes192downloads

NHLOCAL/project-ben-yehuda · main · files are served by the source, never re-hosted here