CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cyttic /trocr-hebrew-syntheticimage100K<n<1M0 likes494 downloads4mo agoHugging Face02cyttic /trocr-hebrew-synthetic-cleanimage100K<n<1M0 likes303 downloads3mo agoHugging Face03cyttic /diffusionpen-hebrew-handwriting DiffusionPen Hebrew Handwriting A large synthetic dataset of Hebrew handwritten text lines with ground-truth transcriptions, for training and evaluating handwritten text recognition (HTR / OCR) models. Every image is a single line of right-to-left Hebrew handwriting synthesized by DiffusionPen — a style-conditioned latent-diffusion handwriting generator — in one of 491 distinct writer styles, and quality-filtered by an independent OCR pass. 149,952 line images, 491 writer… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting.imageimage-to-text100K<n<1M1 likes284 downloads3mo agoHugging Face04cyttic /trocr-hebrew-freefonts-BYimage100K<n<1M0 likes175 downloads22d agoHugging Face05cyttic /trocr-hebrew-synthetic-modernimage100K<n<1M0 likes160 downloads3mo agoHugging Face06community-datasets /hebrew_this_world Dataset Card for HebrewSentiment Dataset Summary HebrewThisWorld is a data set consists of 2028 issues of the newspaper 'This World' edited by Uri Avnery and were published between 1950 and 1989. Released under the AGPLv3 license. Data Annotation: Supported Tasks and Leaderboards Language modeling Languages Hebrew Dataset Structure csv file with "," delimeter Data Instances Sample: { "issue_num": 637, "page_count": 16… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hebrew_this_world.imagetext-generation1K<n<10K1 likes147 downloads2y agoHugging Face07ivrit-ai /hebrew-handwriting-ocr-benchmarkgated Hebrew Handwriting OCR Benchmark A small, human-verified benchmark for OCR / handwritten text recognition (HTR) on modern Hebrew handwriting: 225 gold lines across 10 pages, one page per writer, drawn from the transcriptor.ivrit.ai volunteer transcription corpus. This is a test set. There is no train split, by design — it exists to be held out. It is deliberately small and clean rather than large and noisy: every line was transcribed by at least two volunteers independently and… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/hebrew-handwriting-ocr-benchmark.imageimage-to-textn<1K0 likes147 downloads19d agoHugging Face08cyttic /diffusionpen-hebrew-handwriting-cer0 DiffusionPen Hebrew Handwriting — CER=0 clean subset The highest-fidelity slice of DiffusionPen Hebrew Handwriting: only the 33,082 line images that an independent Hebrew TrOCR read back with an exact match (character error rate = 0.0). Built to test whether a smaller, label-clean set trains a better recognizer than the full (noisier) 150k set. 33,082 line images, 491 writer styles Splits (writer-independent, style-disjoint): train 26,607 / validation 3,187 / test 3,288 Every… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/diffusionpen-hebrew-handwriting-cer0.imageimage-to-text10K<n<100K0 likes129 downloads3mo agoHugging Face09samaritan-ai /hebrew_synth_linesimagetext-generation100K<n<1M1 likes117 downloads1y agoHugging Face10cyttic /trocr-hebrew-freefonts9image100K<n<1M0 likes117 downloads1mo agoHugging Face11cyttic /trocr-hebrew-matanimage1K<n<10K0 likes99 downloads3mo agoHugging Face12mr3vial /paleo-hebrew-seals-synthetic PaleoHebrew-Seals Synthetic Corpus This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions. Why this dataset is needed Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level. Overview The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.imageobject-detection100K<n<1M0 likes93 downloads4mo agoHugging Face13johnlockejrr /hebrew_synthimage10K<n<100K0 likes73 downloads1y agoHugging Face14johnlockejrr /hebrew_synth_trocrimage100K<n<1M0 likes67 downloads1y agoHugging Face15mr3vial /paleo-hebrew-seals-unambiguous PaleoHebrew-Seals Real Benchmark (Unambiguous Subset) This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs. Why this dataset is needed Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.imageobject-detectionn<1K0 likes64 downloads6mo agoHugging Face16isaacmg /synthetic_hebrew_v3 Synthetic Multi-Font Hebrew OCR v3 Part of the Cairo Genizah AI Project 12,000 synthetic images of unmemorizable Hebrew text across 12 typefaces — the font-generalization successor to synthetic_rashi. Same design principle: the text cannot be recited from a language prior (shuffled corpus words, random character strings, confusable-letter drills), so success requires reading glyphs. Font weighting follows a measured 20-typeface probe of where fine-tuned Hebrew VLMs actually… See the full description on the dataset page: https://huggingface.co/datasets/isaacmg/synthetic_hebrew_v3.imageimage-to-text10K<n<100K0 likes62 downloads24d agoHugging Face17sivan22 /hebrew-handwritten-dataset Dataset Information Keywords Hebrew, handwritten, letters Description HDD_v0 consists of images of isolated Hebrew characters together with training and test sets subdivision. The images were collected from hand-filled forms. For more details, please refer to [1]. When using this dataset in research work, please cite [1]. [1] I. Rabaev, B. Kurar Barakat, A. Churkin and J. El-Sana. The HHD Dataset. The 17th International Conference on Frontiers in Handwriting… See the full description on the dataset page: https://huggingface.co/datasets/sivan22/hebrew-handwritten-dataset.imageimage-classification1K<n<10K15 likes54 downloads3y agoHugging Face18johnlockejrr /hebrew_synth_linesimage10K<n<100K0 likes37 downloads1y agoHugging Face19youssefkhalil320 /hebrew-ocr-doctags-dataset_v2image10K<n<100K0 likes30 downloads1y agoHugging Face20danielrosehill /Hebrew-Language-Signage Hebrew Language Signage Dataset Overview This dataset contains photographs of Hebrew language text in everyday contexts throughout Israel, with a particular focus on signage displays including street signs, commercial signage, and public information displays. Dataset Details Total Images: 68 Format: PNG Content: Real-world photographs of Hebrew text and signage Language Coverage: Primarily Hebrew, with many signs also containing English and Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Hebrew-Language-Signage.imagen<1K0 likes28 downloads11mo agoHugging Face21cyttic /trocr-hebrew-humanimage1K<n<10K0 likes26 downloads4mo agoHugging Face22AmitKabya /hebrew-doc-ocr-benchmarkimagen<1K0 likes17 downloads5mo agoHugging Face23sivan22 /hebrew-handwritten-characters Dataset Information Keywords Hebrew, handwritten, letters Description HDD_v0 consists of images of isolated Hebrew characters together with training and test sets subdivision. The images were collected from hand-filled forms. For more details, please refer to [1]. When using this dataset in research work, please cite [1]. [1] I. Rabaev, B. Kurar Barakat, A. Churkin and J. El-Sana. The HHD Dataset. The 17th International Conference on Frontiers in Handwriting… See the full description on the dataset page: https://huggingface.co/datasets/sivan22/hebrew-handwritten-characters.image1K<n<10K1 likes12 downloads3y agoHugging Face24sivan22 /hebrew-words-dataset Dataset Card for "hebrew-words-dataset" More Information needed imagen<1K0 likes8 downloads3y agoHugging Face25youssefkhalil320 /hebrew_images_doc_tags_all_v6image1K<n<10K0 likes8 downloads1y agoHugging Face26youssefkhalil320 /hebrew_images_doc_tags_all_v15image10K<n<100K0 likes8 downloads11mo agoHugging Face27youssefkhalil320 /hebrew-ocr-doctags-datasetimage1K<n<10K0 likes6 downloads1y agoHugging Face28youssefkhalil320 /hebrew_images_doc_tags_all_v12image10K<n<100K0 likes6 downloads11mo agoHugging Face29youssefkhalil320 /hebrew_images_doc_tags_all_v5image1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.