CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cyttic /heb-words27-spaced heb-words27-diffpen Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit is a whole word. Content 20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated Hebrew sentences, 167,879 unique words) 643 numbers -- LLMGen2 contains no digits at all, because the generation prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-spaced.imageimage-to-text100K<n<1M0 likes571 downloads4d agoHugging Face02icfoss /malayalam-ocr-words Malayalam OCR Words A word-level Malayalam OCR dataset: cropped word images paired with their transcribed text label, split into train/validation/test sets. Dataset structure train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word> train/ val/ test/ # image files referenced by the corresponding CSV Each CSV row maps one image file (path relative to its split folder) to its ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.imageimage-to-text10K<n<100K0 likes340 downloads8d agoHugging Face03shuangzhiaishang /Chinese-Char-Wordsimage1K<n<10K0 likes319 downloads1y agoHugging Face04grascii /gregg-anniversary-words Gregg Anniversary Words This dataset is derived from the 1930 Gregg Shorthand Dictionary1. Structure The dataset contains three columns: image: The image of a shorthand form in the dictionary grascii_normalized: The normalized grascii of the shorthand form longhand: The English longhand represented by the shorthand form Issues If you notice any problems in the dataset, open an issue in the datasets repository. 1Gregg, John Robert. Gregg Shorthand Dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-anniversary-words.imageimage-to-text10K<n<100K0 likes276 downloads5mo agoHugging Face05alix-tz /noisy-gt-missing-words Noisy Ground Truth - Missing Words Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator. In Noisy Ground Truth - Missing Words, each variation column is affected by the noise, without considering the split between train, validation and test. Data structure The dataset is composed of the… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words.imageimage-to-text10K<n<100K0 likes263 downloads1y agoHugging Face06c3rl /IIIT-INDIC-HW-WORDS-Hindi IIIT-INDIC-HW-WORDS-Hindi Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images. Overview The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.imageimage-to-text10K<n<100K5 likes244 downloads2y agoHugging Face07sin10122484 /indonesian_words_imageimage1K<n<10K0 likes222 downloads20d agoHugging Face08cyttic /heb-words27-diffpen heb-words27-diffpen Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit is a whole word. Content 20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated Hebrew sentences, 167,879 unique words) 643 numbers -- LLMGen2 contains no digits at all, because the generation prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-diffpen.imageimage-to-text100K<n<1M0 likes199 downloads5d agoHugging Face09alix-tz /noisy-gt-missing-words-train-only Noisy Ground Truth - Missing Words in Train Split only Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator. In Noisy Ground Truth - Missing Words in Train Split only, each variation column is affected by the noise, only when the split is for training. The validation and test splits are not affected by… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words-train-only.imageimage-to-text1K<n<10K0 likes197 downloads2y agoHugging Face10priyank-m /IAM_words_text_recognitionimage100K<n<1M9 likes190 downloads4y agoHugging Face11sin10122484 /iam-words-no_symbolsimage1K<n<10K0 likes177 downloads24d agoHugging Face12alix-tz /noisy-gt-xxx-words Noisy Ground Truth - Words Replaced with XXX Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator. In Noisy Ground Truth - Words Replaced with XXX, each variation column is affected by the noise, without considering the split between train, validation and test. Data structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-xxx-words.imageimage-to-text1K<n<10K0 likes128 downloads2y agoHugging Face13c3rl /IIIT-INDIC-HW-WORDS-Tamil IIIT-INDIC-HW-WORDS-Tamil Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images. Overview The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Tamil words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Tamil.image100K<n<1M4 likes118 downloads2y agoHugging Face14biglam /loc_beyond_words Dataset Card for Beyond Words Dataset Summary The Beyond Words dataset is a crowdsourced collection of bounding box annotations on World War I-era historical newspaper pages from the Library of Congress’s Chronicling America collection. Volunteers marked seven types of visual content — photographs, illustrations, maps, comics, editorial cartoons, headlines, and advertisements — enabling the training of the visual content recognition model behind the Newspaper Navigator… See the full description on the dataset page: https://huggingface.co/datasets/biglam/loc_beyond_words.imageobject-detection1K<n<10K15 likes107 downloads1y agoHugging Face15Rapidata /sora-video-generation-aligned-words Rapidata Video Generation Word for Word Alignment Dataset If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~1500 human evaluators were asked to evaluate AI-generated videos based on what part of the prompt did not align the video. The specific… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-aligned-words.imagevideo-classificationn<1K19 likes88 downloads2y agoHugging Face16SARAH-HADDAD /Amount_in_Arabic_words_dzimagen<1K0 likes84 downloads1y agoHugging Face17priyank-m /trdg_random_single_words_en_text_recognition Dataset Card for "trdg_random_single_words_en_text_recognition" More Information needed image100K<n<1M0 likes74 downloads4y agoHugging Face18alix-tz /noisy-gt-xxx-words-train-only Noisy Ground Truth - Words Replaced with XXX in Train Split only Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator. In Noisy Ground Truth - Words Replaced with XXX in Train Split only, each variation column is affected by the noise, without considering the split between train, validation and test.… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-xxx-words-train-only.imageimage-to-text1K<n<10K0 likes61 downloads2y agoHugging Face19avi-kai /Medical_Prescription_Handwritten_Words Medical Prescription Handwritten Words This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain. Structure images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.) data.csv: Maps each image file to its corresponding label (word) Example Use Cases OCR (Optical… See the full description on the dataset page: https://huggingface.co/datasets/avi-kai/Medical_Prescription_Handwritten_Words.imageimage-classificationn<1K1 likes50 downloads1y agoHugging Face20priyank-m /trdg_dict_random_words_en_text_recognition Dataset Card for "trdg_random_words_en_text_recognition" More Information needed image100K<n<1M0 likes42 downloads4y agoHugging Face21MMMuzammil /Medical_Prescription_Handwritten_Words Medical Prescription Handwritten Words This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain. Structure images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.) data.csv: Maps each image file to its corresponding label (word) Example Use Cases OCR… See the full description on the dataset page: https://huggingface.co/datasets/MMMuzammil/Medical_Prescription_Handwritten_Words.imageimage-classificationn<1K0 likes31 downloads3mo agoHugging Face22sunildkumar /message-decoding-words-and-sequences-r1image10K<n<100K0 likes28 downloads2y agoHugging Face23harness-race /loc_beyond_words_cocoimage1K<n<10K0 likes28 downloads2mo agoHugging Face24Mouwiya /image-in-Words400 Mouwiya/image-in-words400 Dataset Description Mouwiya/image-in-words400 is a dataset consisting of 400 images along with their corresponding descriptive captions. The dataset is designed for tasks related to image captioning, where the goal is to generate accurate and contextually relevant descriptions for visual content. This dataset can be used to train and evaluate models that bridge the gap between visual and textual data. Dataset Details Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/image-in-Words400.imageimage-to-textn<1K0 likes25 downloads2y agoHugging Face25HarishBonu /IIIT-INDIC-HW-WORDS-Hindi IIIT-INDIC-HW-WORDS-Hindi Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images. Overview The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research… See the full description on the dataset page: https://huggingface.co/datasets/HarishBonu/IIIT-INDIC-HW-WORDS-Hindi.imageimage-to-text10K<n<100K0 likes24 downloads1mo agoHugging Face26grascii /gregg-preanniversary-words Gregg Preanniversary Words This dataset is derived from the 1916 Gregg Shorthand Dictionary1 and builds on top of the Gregg1916 dataset by: Correcting the labels of 250+ images Adding 550+ images for words missing in the original dataset Redoing 100+ poorly cropped images or images with stray marks Structure The dataset contains three columns: image: The image of a shorthand form in the dictionary grascii_normalized: The normalized grascii of the shorthand form… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-preanniversary-words.imageimage-to-text10K<n<100K0 likes21 downloads7mo agoHugging Face27AlhitawiMohammed22 /words_hu_dictThis dataset was generated from a Hungarian dictionary, where 60345 sample given The command used to generate data : python3 run.py -i "dicts/hu.txt" -t 8 -f 64 -l hu -c 60345 -na 2 --output_dir "out/words/hu/" --font_dir fonts/hu/ -b 3 -al 0 TRDGHuMu is used for generating text: https://github.com/Mohammed20201991/TextRecognitionDataGeneratorHuMu23 imageimage-to-text10K<n<100K0 likes19 downloads3y agoHugging Face28Groundlight /message-decoding-words-and-sequences-target-zoom-in-r1image10K<n<100K0 likes18 downloads1y agoHugging Face29Groundlight /message-decoding-words-and-sequences-target-zoom-inimage10K<n<100K0 likes15 downloads1y agoHugging Face30vklinhhh /vnondb_sentence_by_wordsimage10K<n<100K0 likes14 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.