datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trocr-hebrew-synthetictrocr-hebrew-synthetic-cleandiffusionpen-hebrew-handwriting
DiffusionPen Hebrew Handwriting
A large synthetic dataset of Hebrew handwritten text lines with ground-truth
transcriptions, for training and evaluating handwritten text recognition (HTR / OCR)
models. Every image is a single line of right-to-left Hebrew handwriting synthesized by
DiffusionPen — a style-conditioned latent-diffusion handwriting generator — in one of
491 distinct writer styles, and quality-filtered by an independent OCR pass.
149,952 line images, 491 writer… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting.trocr-hebrew-freefonts-BYtrocr-hebrew-synthetic-modernhebrew_this_world
Dataset Card for HebrewSentiment
Dataset Summary
HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license.
Data Annotation:
Supported Tasks and Leaderboards
Language modeling
Languages
Hebrew
Dataset Structure
csv file with "," delimeter
Data Instances
Sample:
{
"issue_num": 637,
"page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.hebrew-handwriting-ocr-benchmark
Hebrew Handwriting OCR Benchmark
A small, human-verified benchmark for OCR / handwritten text recognition (HTR) on
modern Hebrew handwriting: 225 gold lines across 10 pages, one page per
writer, drawn from the transcriptor.ivrit.ai
volunteer transcription corpus.
This is a test set. There is no train split, by design — it exists to be held
out. It is deliberately small and clean rather than large and noisy: every line
was transcribed by at least two volunteers independently and… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/hebrew-handwriting-ocr-benchmark.diffusionpen-hebrew-handwriting-cer0
DiffusionPen Hebrew Handwriting — CER=0 clean subset
The highest-fidelity slice of DiffusionPen Hebrew Handwriting:
only the 33,082 line images that an independent Hebrew TrOCR read back with an exact
match (character error rate = 0.0). Built to test whether a smaller, label-clean set
trains a better recognizer than the full (noisier) 150k set.
33,082 line images, 491 writer styles
Splits (writer-independent, style-disjoint): train 26,607 / validation 3,187 / test 3,288
Every… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting-cer0.hebrew_synth_linestrocr-hebrew-freefonts9trocr-hebrew-matanpaleo-hebrew-seals-synthetic
PaleoHebrew-Seals Synthetic Corpus
This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions.
Why this dataset is needed
Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level.
Overview
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.hebrew_synthhebrew_synth_trocrpaleo-hebrew-seals-unambiguous
PaleoHebrew-Seals Real Benchmark (Unambiguous Subset)
This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs.
Why this dataset is needed
Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.synthetic_hebrew_v3
Synthetic Multi-Font Hebrew OCR v3
Part of the Cairo Genizah AI Project
12,000 synthetic images of unmemorizable Hebrew text across 12
typefaces — the font-generalization successor to
synthetic_rashi.
Same design principle: the text cannot be recited from a language prior
(shuffled corpus words, random character strings, confusable-letter drills),
so success requires reading glyphs.
Font weighting follows a measured 20-typeface probe of where fine-tuned
Hebrew VLMs actually… See the full description on the dataset page: https://huggingface.co/datasets/isaacmg/synthetic_hebrew_v3.hebrew-handwritten-dataset
Dataset Information
Keywords
Hebrew, handwritten, letters
Description
HDD_v0 consists of images of isolated Hebrew characters together with training and test sets subdivision.
The images were collected from hand-filled forms.
For more details, please refer to [1].
When using this dataset in research work, please cite [1].
[1] I. Rabaev, B. Kurar Barakat, A. Churkin and J. El-Sana. The HHD Dataset. The 17th International Conference on Frontiers in Handwriting… See the full description on the dataset page: https://huggingface.co/datasets/sivan22/hebrew-handwritten-dataset.hebrew_synth_lineshebrew-ocr-doctags-dataset_v2Hebrew-Language-Signage
Hebrew Language Signage Dataset
Overview
This dataset contains photographs of Hebrew language text in everyday contexts throughout Israel, with a particular focus on signage displays including street signs, commercial signage, and public information displays.
Dataset Details
Total Images: 68
Format: PNG
Content: Real-world photographs of Hebrew text and signage
Language Coverage: Primarily Hebrew, with many signs also containing English and Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Hebrew-Language-Signage.trocr-hebrew-humanhebrew-doc-ocr-benchmarkhebrew-handwritten-characters
Dataset Information
Keywords
Hebrew, handwritten, letters
Description
HDD_v0 consists of images of isolated Hebrew characters together with training and test sets subdivision.
The images were collected from hand-filled forms.
For more details, please refer to [1].
When using this dataset in research work, please cite [1].
[1] I. Rabaev, B. Kurar Barakat, A. Churkin and J. El-Sana. The HHD Dataset. The 17th International Conference on Frontiers in Handwriting… See the full description on the dataset page: https://huggingface.co/datasets/sivan22/hebrew-handwritten-characters.hebrew-words-dataset
Dataset Card for "hebrew-words-dataset"
More Information needed
hebrew_images_doc_tags_all_v6hebrew_images_doc_tags_all_v15hebrew-ocr-doctags-datasethebrew_images_doc_tags_all_v12hebrew_images_doc_tags_all_v5
