CoolFace
Datasetpublic

mbole/hebrew-tzfira-dataset

Ha-Tsfira OCR and POS-Tagged Dataset Dataset Summary This dataset contains OCR-processed and POS-tagged text from Ha-Tsfira (הצפירה), a Hebrew-language newspaper published in Poland from 1862 and then from 1874 to 1931. The dataset includes 50 newspaper issues that have been digitized, cleaned, and linguistically annotated. Languages Hebrew (he) Dataset Structure DatasetDict({ train: Dataset({ features: ['id', 'ocr_text'… See the full description on the dataset page: https://huggingface.co/datasets/mbole/hebrew-tzfira-dataset.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes9downloads

mbole/hebrew-tzfira-dataset · main · files are served by the source, never re-hosted here