CoolFace
Datasetpublic

tomron87/hebrew-wikipedia-sentences-corpus

Hebrew Wikipedia Sentences Corpus A corpus of 10,999,257 cleaned, deduplicated Hebrew sentences extracted from 366,610 Hebrew Wikipedia articles. Dataset Description This dataset contains Hebrew sentences extracted from Hebrew Wikipedia (crawled 2026-02). Each sentence has been cleaned, filtered for quality, and deduplicated. The dataset is intended for Hebrew NLP tasks including language modeling, text classification, NER, sentence similarity, and more.… See the full description on the dataset page: https://huggingface.co/datasets/tomron87/hebrew-wikipedia-sentences-corpus.

sourceHugging Facecc-by-sa-3.0updated 8mo agoView on Hugging Face
0likes48downloads
3 commits on main
76fe89e8mo ago

Upload README.md with huggingface_hub

tomron87
fc442ff8mo ago

Upload sentences.parquet with huggingface_hub

tomron87
5a85a3b8mo ago

initial commit

tomron87