CoolFace
Datasetpublic

tomron87/hebrew-wikipedia-sentences-corpus

Hebrew Wikipedia Sentences Corpus A corpus of 10,999,257 cleaned, deduplicated Hebrew sentences extracted from 366,610 Hebrew Wikipedia articles. Dataset Description This dataset contains Hebrew sentences extracted from Hebrew Wikipedia (crawled 2026-02). Each sentence has been cleaned, filtered for quality, and deduplicated. The dataset is intended for Hebrew NLP tasks including language modeling, text classification, NER, sentence similarity, and more.… See the full description on the dataset page: https://huggingface.co/datasets/tomron87/hebrew-wikipedia-sentences-corpus.

sourceHugging Facecc-by-sa-3.0updated 8mo agoView on Hugging Face
0likes48downloads
settings

This repository belongs to tomron87 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehebrew-wikipedia-sentences-corpus
visibilitypublic
licencecc-by-sa-3.0
gatedno
ownertomron87
Account settings
tomron87/hebrew-wikipedia-sentences-corpus · CoolFace